paper-with-me

Papers

Emotional Prosody Control for Speech Generation

2021-11-07 · Sarath Sivaprasad, Saiteja Kosgi, Vineet Gandhi

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from prosody sequences in training data or transferred from a source style. We propose a text to speech(TTS) system, where a user can choose the emotion of generated speech from a continuous and meaningful emotion space (Arousal-Valence space). The proposed TTS system can generate speech from the text in any speaker's style, with fine control of emotion. We show that the system works on emotion unseen during training and can scale to previously unseen speakers given his/her speech sample. Our work expands the horizon of the state-of-the-art FastSpeech2 backbone to a multi-speaker setting and gives it much-coveted continuous (and interpretable) affective control, without any observable degradation in the quality of the synthesized speech.

📄 PDF Abstract BibTeX arXiv:2111.04730

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

2025-08-09 · Jinsung Yoon, Wooyeol Jeong, Jio Gim, Young-Joo Suh arxiv

Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotiona…

Voice ConversionSpeech Synthesis

Robust and fine-grained prosody control of end-to-end speech synthesis

2018-11-06 · Young-Gun Lee, Taesu Kim

We propose prosody embeddings for emotional and expressive speech synthesis networks. The proposed methods introduce temporal structures in the embedding networks, thus enabling fine-grained control of the speaking style…

Expressive Speech SynthesisSpeech Synthesis

Bridging the prosody GAP: Genetic Algorithm with People to efficiently sample emotional prosody

2022-05-10 · Pol van Rijn, Harin Lee, Nori Jacoby

The human voice effectively communicates a range of emotions with nuanced variations in acoustics. Existing emotional speech corpora are limited in that they are either (a) highly curated to induce specific emotions with…

PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts

2025-05-27 · Tianhua Qi, Shiyan Wang, Cheng Lu, Tengfei Song 외

Controllable emotional voice conversion (EVC) aims to manipulate emotional expressions to increase the diversity of synthesized speech. Existing methods typically rely on predefined labels, reference audios, or prespecif…

DiversityRhythmVoice Conversion

Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2

2026-03-12 · Suvendu Sekhar Mohanty arxiv

We propose a novel causal prosody mediation framework for expressive text-to-speech (TTS) synthesis. Our approach augments the FastSpeech2 architecture with explicit emotion conditioning and introduces counterfactual tra…

Speech Synthesis