paper-with-me

홈 › Papers

Adaptive Duration Model for Text Speech Alignment

2025-07-30 · Junjie Cao arxiv

Speech-to-text alignment is a critical component of neural text to speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-line, while non-autoregressive end to end TTS models rely on durations extracted from external sources. In this paper, we propose a novel duration prediction framework that can give promising phoneme-level duration distribution with given text. In our experiments, the proposed duration model has more precise prediction and adaptation ability to conditions, compared to previous baseline models. Specifically, it makes a considerable improvement on phoneme-level alignment accuracy and makes the performance of zero-shot TTS models more robust to the mismatch between prompt audio and input audio.

📄 PDF Abstract BibTeX arXiv:2507.22612

Code (0)

등록된 구현이 없습니다.

Tasks

Text to Speech

Similar Papers 제목 키워드 기반

AutoTTS: End-to-End Text-to-Speech Synthesis through Differentiable Duration Modeling

2022-03-21 · Bac Nguyen, Fabien Cardinaux, Stefan Uhlich

Parallel text-to-speech (TTS) models have recently enabled fast and highly-natural speech synthesis. However, they typically require external alignment models, which are not necessarily optimized for the decoder as they …

DecoderSpeech Synthesistext-to-speechText to Speech+1

Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization

2025-08-12 · Chaoqun Cui, Liangbin Huang, Shijing Wang, Zhe Tong 외 arxiv

Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information …

Machine Translation

JDI-T: Jointly trained Duration Informed Transformer for Text-To-Speech without Explicit Alignment

2020-05-15 · Dan Lim, Won Jang, Gyeonghwan O, Heayoung Park 외

We propose Jointly trained Duration Informed Transformer (JDI-T), a feed-forward Transformer with a duration predictor jointly trained without explicit alignments in order to generate an acoustic feature sequence from an…

text-to-speechText to Speech

DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment

2024-01-16 · Hyung-Seok Oh, Sang-Hoon Lee, Deok-Hyeon Cho, Seong-Whan Lee

Emotional voice conversion (EVC) involves modifying various acoustic characteristics, such as pitch and spectral envelope, to match a desired emotional state while preserving the speaker's identity. Existing EVC methods …

DisentanglementSelf-Supervised LearningText to SpeechVoice Conversion

DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis

2023-09-22 · Yu Gu, Yianrao Bian, Guangzhi Lei, Chao Weng 외

This paper introduces an improved duration informed attention neural network (DurIAN-E) for expressive and high-fidelity text-to-speech (TTS) synthesis. Inherited from the original DurIAN model, an auto-regressive model …

DenoisingSpeech Synthesistext-to-speechText to Speech+1