Adaptive Duration Model for Text Speech Alignment
Speech-to-text alignment is a critical component of neural text to speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-line, while non-autoregressive end to end TTS models rely on durations extracted from external sources. In this paper, we propose a novel duration prediction framework that can give promising phoneme-level duration distribution with given text. In our experiments, the proposed duration model has more precise prediction and adaptation ability to conditions, compared to previous baseline models. Specifically, it makes a considerable improvement on phoneme-level alignment accuracy and makes the performance of zero-shot TTS models more robust to the mismatch between prompt audio and input audio.
Code (0)
등록된 구현이 없습니다.
Tasks
Text to SpeechSimilar Papers 제목 키워드 기반
AutoTTS: End-to-End Text-to-Speech Synthesis through Differentiable Duration Modeling
Parallel text-to-speech (TTS) models have recently enabled fast and highly-natural speech synthesis. However, they typically require external alignment models, which are not necessarily optimized for the decoder as they …
DecoderSpeech Synthesistext-to-speechText to Speech+1Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information …
Machine TranslationJDI-T: Jointly trained Duration Informed Transformer for Text-To-Speech without Explicit Alignment
We propose Jointly trained Duration Informed Transformer (JDI-T), a feed-forward Transformer with a duration predictor jointly trained without explicit alignments in order to generate an acoustic feature sequence from an…
text-to-speechText to SpeechDurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
Emotional voice conversion (EVC) involves modifying various acoustic characteristics, such as pitch and spectral envelope, to match a desired emotional state while preserving the speaker's identity. Existing EVC methods …
DisentanglementSelf-Supervised LearningText to SpeechVoice ConversionDurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis
This paper introduces an improved duration informed attention neural network (DurIAN-E) for expressive and high-fidelity text-to-speech (TTS) synthesis. Inherited from the original DurIAN model, an auto-regressive model …
DenoisingSpeech Synthesistext-to-speechText to Speech+1