paper-with-me

홈 › Papers

Attentron: Few-Shot Text-to-Speech Utilizing Attention-Based Variable-Length Embedding

2020-08-12

On account of growing demands for personalization, the need for a so-called few-shot TTS system that clones speakers with only a few data is emerging. To address this issue, we propose Attentron, a few-shot TTS model that clones voices of speakers unseen during training. It introduces two special encoders, each serving different purposes. A fine-grained encoder extracts variable-length style information via an attention mechanism, and a coarse-grained encoder greatly stabilizes the speech synthesis, circumventing unintelligible gibberish even for synthesizing speech of unseen speakers. In addition, the model can scale out to an arbitrary number of reference audios to improve the quality of the synthesized speech. According to our experiments, including a human evaluation, the proposed model significantly outperforms state-of-the-art models when generating speech for unseen speakers in terms of speaker similarity and quality.

📄 PDF Abstract BibTeX arXiv:2005.08484

Code (1)

jasminsternkopf/mel_cepstral_distance

Tasks

Speech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

2024-07-07 · Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu 외

Recent years have witnessed a trend that large language model (LLM) based text-to-speech (TTS) emerges into the mainstream due to their high naturalness and zero-shot capacity. In this paradigm, speech signals are discre…

Language ModellingLarge Language ModelQuantizationspeech-recognition+5

Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment

2024-06-25 · Paarth Neekhara, Shehzeen Hussain, Subhankar Ghosh, Jason Li 외

Large Language Model (LLM) based text-to-speech (TTS) systems have demonstrated remarkable capabilities in handling large speech datasets and generating natural speech for new speakers. However, LLM-based TTS models are …

DecoderLanguage ModelingLanguage ModellingLarge Language Model+3

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

2025-06-03 · Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse…

Expressive Speech SynthesisPrompt LearningSpeech Synthesistext-to-speech+1

MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech

2024-10-04 · Taejun Bak, Youngsik Eom, SeungJae Choi, Young-Sun Joo

Text-to-speech (TTS) systems that scale up the amount of training data have achieved significant improvements in zero-shot speech synthesis. However, these systems have certain limitations: they require a large amount of…

DisentanglementSpeech SynthesisStyle Transfertext-to-speech+1

Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders

2025-02-18 · Seungbae Kim, Daeun Lee, Brielle Stark, Jinyoung Han

Individuals with language disorders often face significant communication challenges due to their limited language processing and comprehension abilities, which also affect their interactions with voice-assisted systems t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+5