PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
Controllable emotional voice conversion (EVC) aims to manipulate emotional expressions to increase the diversity of synthesized speech. Existing methods typically rely on predefined labels, reference audios, or prespecified factor values, often overlooking individual differences in emotion perception and expression. In this paper, we introduce PromptEVC that utilizes natural language prompts for precise and flexible emotion control. To bridge text descriptions with emotional speech, we propose emotion descriptor and prompt mapper to generate fine-grained emotion embeddings, trained jointly with reference embeddings. To enhance naturalness, we present a prosody modeling and control pipeline that adjusts the rhythm based on linguistic content and emotional cues. Additionally, a speaker encoder is incorporated to preserve identity. Experimental results demonstrate that PromptEVC outperforms state-of-the-art controllable EVC methods in emotion conversion, intensity control, mixed emotion synthesis, and prosody manipulation. Speech samples are available at https://jeremychee4.github.io/PromptEVC/.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityRhythmVoice ConversionSimilar Papers 제목 키워드 기반
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
Realistic emotional voice conversion (EVC) aims to enhance emotional diversity of converted audios, making the synthesized voices more authentic and natural. To this end, we propose Emotional Intensity-aware Network (EIN…
DiversityRhythmVoice ConversionMaestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotiona…
Voice ConversionSpeech SynthesisSeeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
Face-based Voice Conversion (FVC) is a novel task that leverages facial images to generate the target speaker's voice style. Previous work has two shortcomings: (1) suffering from obtaining facial embeddings that are wel…
Contrastive LearningDisentanglementDiversityVoice ConversionZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism
The human voice conveys not just words but also emotional states and individuality. Emotional voice conversion (EVC) modifies emotional expressions while preserving linguistic content and speaker identity, improving appl…
Emotion ClassificationVoice ConversionStarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion
We present an unsupervised non-parallel many-to-many voice conversion (VC) method using a generative adversarial network (GAN) called StarGAN v2. Using a combination of adversarial source classifier loss and perceptual l…
Generative Adversarial Networktext-to-speechText to SpeechVoice Conversion