Trích chọn các tham số đặc trưng tiếng nói cho hệ thống tổng hợp tiếng Việt dựa vào mô hình Markov ẩn

The overall performance of the systems is often limited by the accuracy of the underlying speech parameterization and reconstruction method. The method proposed in this paper allows accurate MFCC, F0 and tone extraction and high-quality reconstruction of speech signals assuming Mel Log Spectral Approximation filter. Its suitability for high-quality HMM-based speech synthesis is shown through evaluations subjectively.

Từ khóa: Vietnamese speech synthesis, Context-dependent, Speech parameterization, Statistical parametric speech synthesis, Hidden Markov Models, Mel-frequency cepstral coefficient, Fundamental frequency