Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting the speech rate on speech quality, such as intelligibility and speaker characteristics, has been underexplored. In this work, we propose a novel total-duration-aware (TDA) duration model for TTS, where phoneme durations are predicted not only from the text input but also from an additional input of the total target duration. We also propose a MaskGIT-based duration model that enhances the diversity and quality of the predicted phoneme durations. Our results demonstrate that the proposed TDA duration models achieve better intelligibility and speaker similarity for various speech rate configurations compared to the baseline models. We also show that the proposed MaskGIT-based model can generate phoneme durations with higher quality and diversity compared to its regression or flow-matching counterparts.
Code (0)
등록된 구현이 없습니다.
Tasks
Diversitytext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
AutoTTS: End-to-End Text-to-Speech Synthesis through Differentiable Duration Modeling
Parallel text-to-speech (TTS) models have recently enabled fast and highly-natural speech synthesis. However, they typically require external alignment models, which are not necessarily optimized for the decoder as they …
DecoderSpeech Synthesistext-to-speechText to Speech+1End-to-End Text-to-Speech using Latent Duration based on VQ-VAE
Explicit duration modeling is a key to achieving robust and efficient alignment in text-to-speech synthesis (TTS). We propose a new TTS framework using explicit duration modeling that incorporates duration as a discrete …
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisAn experimental and computational study of an Estonian single-person word naming
This study investigates lexical processing in Estonian. A large-scale single-subject experiment is reported that combines the word naming task with eye-tracking. Five response variables (first fixation duration, total fi…
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. In this paper, we propose a "source-synth…
Using Gaze to Predict Text Readability
We show that text readability prediction improves significantly from hard parameter sharing with models predicting first pass duration, total fixation duration and regression duration. Specifically, we induce multi-task …
Machine TranslationMulti-Task LearningregressionSentence+2