Maximum Voiced Frequency Estimation: Exploiting Amplitude and Phase Spectra
Maximum Voiced Frequency (MVF) is used in various speech models as the spectral boundary separating periodic and aperiodic components during the production of voiced sounds. Recent studies have shown that its proper estimation and modeling enhance the quality of statistical parametric speech synthesizers. Contrastingly, these same methods of MVF estimation have been reported to degrade the performance of singing voice synthesizers. This paper proposes a new approach for MVF estimation which exploits both amplitude and phase spectra. It is shown that phase conveys relevant information about the harmonicity of the voice signal, and that it can be jointly used with features derived from the amplitude spectrum. This information is further integrated into a maximum likelihood criterion which provides a decision about the MVF estimate. The proposed technique is compared to two state-of-the-art methods, and shows a superior performance in both objective and subjective evaluations. Perceptual tests indicate a drastic improvement in high-pitched voices.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A New Statistic Feature of the Short-Time Amplitude Spectrum Values for Human's Unvoiced Pronunciation
In this paper, a new statistic feature of the discrete short-time amplitude spectrum is discovered by experiments for the signals of unvoiced pronunciation. For the random-varying short-time spectrum, this feature reveal…
A Fast and Accurate Pitch Estimation Algorithm Based on the Pseudo Wigner-Ville Distribution
Estimation of fundamental frequency (F0) in voiced segments of speech signals, also known as pitch tracking, plays a crucial role in pitch synchronous speech analysis, speech synthesis, and speech manipulation. In this p…
Speech SynthesisRegression-based Melody Estimation with Uncertainty Quantification
Existing machine learning models approach the task of melody estimation from polyphonic audio as a classification problem by discretizing the pitch values, which results in the loss of finer frequency variations present …
regressionUncertainty QuantificationSpeech Decomposition Based on a Hybrid Speech Model and Optimal Segmentation
In a hybrid speech model, both voiced and unvoiced components can coexist in a segment. Often, the voiced speech is regarded as the deterministic component, and the unvoiced speech and additive noise are the stochastic c…
Model SelectionSegmentationHigh-Resolution Multipath Angle Estimation Based on Power-Angle-Delay Profile for Directional Scanning Sounding
Directional scanning sounding (DSS) has become widely adopted for high-frequency channel measurements because it effectively compensates for severe path loss. However, the resolution of existing multipath component (MPC)…