paper-with-me

Papers

Total-Duration-Aware Duration Modeling for Text-to-Speech Systems

2024-06-06 · Sefik Emre Eskimez, Xiaofei Wang, Manthan Thakker, Chung-Hsien Tsai, Canrun Li, Zhen Xiao, Hemin Yang, Zirun Zhu, Min Tang, Jinyu Li, Sheng Zhao, Naoyuki Kanda

Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting the speech rate on speech quality, such as intelligibility and speaker characteristics, has been underexplored. In this work, we propose a novel total-duration-aware (TDA) duration model for TTS, where phoneme durations are predicted not only from the text input but also from an additional input of the total target duration. We also propose a MaskGIT-based duration model that enhances the diversity and quality of the predicted phoneme durations. Our results demonstrate that the proposed TDA duration models achieve better intelligibility and speaker similarity for various speech rate configurations compared to the baseline models. We also show that the proposed MaskGIT-based model can generate phoneme durations with higher quality and diversity compared to its regression or flow-matching counterparts.

📄 PDF Abstract BibTeX arXiv:2406.04281

Code (0)

등록된 구현이 없습니다.

Tasks

Diversitytext-to-speechText to Speech

Similar Papers 제목 키워드 기반

AutoTTS: End-to-End Text-to-Speech Synthesis through Differentiable Duration Modeling

2022-03-21 · Bac Nguyen, Fabien Cardinaux, Stefan Uhlich

Parallel text-to-speech (TTS) models have recently enabled fast and highly-natural speech synthesis. However, they typically require external alignment models, which are not necessarily optimized for the decoder as they …

DecoderSpeech Synthesistext-to-speechText to Speech+1

End-to-End Text-to-Speech using Latent Duration based on VQ-VAE

2020-10-19 · Yusuke Yasuda, Xin Wang, Junichi Yamagishi

Explicit duration modeling is a key to achieving robust and efficient alignment in text-to-speech synthesis (TTS). We propose a new TTS framework using explicit duration modeling that incorporates duration as a discrete …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

An experimental and computational study of an Estonian single-person word naming

2025-09-03 · Kaidi Lõo, Arvi Tavast, Maria Heitmeier, Harald Baayen arxiv

This study investigates lexical processing in Estonian. A large-scale single-subject experiment is reported that combines the word naming task with eye-tracking. Five response variables (first fixation duration, total fi…

CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data

2026-02-22 · Qibing Bai, Shuhao Shi, Shuai Wang, Yukai Ju 외 arxiv

Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. In this paper, we propose a "source-synth…

Using Gaze to Predict Text Readability

2017-09-01 · WS 2017 9 · Ana Valeria Gonz{\'a}lez-Gardu{\~n}o, Anders S{\o}gaard

We show that text readability prediction improves significantly from hard parameter sharing with models predicting first pass duration, total fixation duration and regression duration. Specifically, we induce multi-task …

Machine TranslationMulti-Task LearningregressionSentence+2