paper-with-me

홈 › Papers

Mixed-EVC: Mixed Emotion Synthesis and Control in Voice Conversion

2022-10-25 · Kun Zhou, Berrak Sisman, Carlos Busso, Bin Ma, Haizhou Li

Emotional voice conversion (EVC) traditionally targets the transformation of spoken utterances from one emotional state to another, with previous research mainly focusing on discrete emotion categories. This paper departs from the norm by introducing a novel perspective: a nuanced rendering of mixed emotions and enhancing control over emotional expression. To achieve this, we propose a novel EVC framework, Mixed-EVC, which only leverages discrete emotion training labels. We construct an attribute vector that encodes the relationships among these discrete emotions, which is predicted using a ranking-based support vector machine and then integrated into a sequence-to-sequence (seq2seq) EVC framework. Mixed-EVC not only learns to characterize the input emotional style but also quantifies its relevance to other emotions during training. As a result, users have the ability to assign these attributes to achieve their desired rendering of mixed emotions. Objective and subjective evaluations confirm the effectiveness of our approach in terms of mixed emotion synthesis and control while surpassing traditional baselines in the conversion of discrete emotions from one to another.

📄 PDF Abstract BibTeX arXiv:2210.13756

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeVoice Conversion

Similar Papers 제목 키워드 기반

Speech Synthesis with Mixed Emotions

2022-08-11 · Kun Zhou, Berrak Sisman, Rajib Rana, B. W. Schuller 외

Emotional speech synthesis aims to synthesize human voices with various emotional effects. The current studies are mostly focused on imitating an averaged style belonging to a specific emotion type. In this paper, we see…

AttributeEmotional Speech SynthesisSpeech Synthesistext-to-speech+1

PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts

2025-05-27 · Tianhua Qi, Shiyan Wang, Cheng Lu, Tengfei Song 외

Controllable emotional voice conversion (EVC) aims to manipulate emotional expressions to increase the diversity of synthesized speech. Existing methods typically rely on predefined labels, reference audios, or prespecif…

DiversityRhythmVoice Conversion

A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

2026-07-01 · Siyi Wang, James Bailey, Ting Dang arxiv

While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood. We present the first comparati…

Speech Synthesis

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

2026-02-03 · Siyi Wang, Shihong Tan, Siyi Liu, Hong Jia 외 arxiv

Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech sys…

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model

2025-09-01 · Joonyong Park, Daisuke Saito, Nobuaki Minematsu arxiv

This study presents a novel approach to voice synthesis that can substitute the traditional grapheme-to-phoneme (G2P) conversion by using a deep learning-based model that generates discrete tokens directly from speech. U…

Self-Supervised LearningSpeech Synthesis