Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
Current emotional text-to-speech systems face challenges in conveying the full spectrum of human emotions, largely due to the inherent complexity of human emotions and the limited range of emotional labels in existing speech datasets. To address these limitations, this paper introduces a TTS framework that provides flexible user control over three emotional dimensions - pleasure, arousal, and dominance - enabling the synthesis of a diverse array of emotional styles. The framework leverages an emotional dimension predictor, trained soley on categorical labels from speech data and grounded in earlier psychological research, which is seamlessly integrated into a language model-based TTS system. Experimental results demonstrates that the proposed framework effectively learns emotional styles from expressive speech, eliminating the need for explicit emotion labels during TTS training, while enhancing the naturalness and diversity of synthesized emotional speech.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeDimensionality ReductionDiversityLanguage ModelingLanguage ModellingSelf-Supervised Learningtext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech
Recent neural codec language models have made great progress in the field of text-to-speech (TTS), but controllable emotional TTS still faces many challenges. Traditional methods rely on predefined discrete emotion label…
Emotional Speech SynthesisLanguage ModelingLanguage ModellingSpeech Synthesis+2EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human …
text-to-speechText to SpeechClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating …
Voice ConversionText-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS
Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the targe…
Language ModelingLanguage ModellingStyle Transfertext-to-speech+1ASR-based Features for Emotion Recognition: A Transfer Learning Approach
During the last decade, the applications of signal processing have drastically improved with deep learning. However areas of affecting computing such as emotional speech synthesis or emotion recognition from spoken langu…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotional Speech SynthesisEmotion Recognition+4