paper-with-me

홈 › Papers

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

2026-02-03 · Siyi Wang, Shihong Tan, Siyi Liu, Hong Jia, Gongping Huang, James Bailey, Ting Dang arxiv

Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech systems enforce a single utterance-level emotion, collapsing affective diversity and suppressing mixed or text-emotion-misaligned expression. While activation steering via latent direction vectors offers a promising solution, it remains unclear whether emotion representations are linearly steerable in TTS, where steering should be applied within hybrid TTS architectures, and how such complex emotion behaviors should be evaluated. This paper presents the first systematic analysis of activation steering for emotional control in hybrid TTS models, introducing a quantitative, controllable steering framework, and multi-rater evaluation protocols that enable composable mixed-emotion synthesis and reliable text-emotion mismatch synthesis. Our results demonstrate, for the first time, that emotional prosody and expressive variability are primarily synthesized by the TTS language module instead of the flow-matching module, and also provide a lightweight steering approach for generating natural, human-like emotional speech.

📄 PDF Abstract BibTeX arXiv:2602.03420

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ZET-Speech: Zero-shot adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models

2023-05-23 · Minki Kang, Wooseok Han, Sung Ju Hwang, Eunho Yang

Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Think Twice: A Human-like Two-stage Conversational Agent for Emotional Response Generation

2023-01-12 · Yushan Qian, Bo wang, Shangzhao Ma, Wu Bin 외

Towards human-like dialogue systems, current emotional dialogue approaches jointly model emotion and semantics with a unified neural network. This strategy tends to generate safe responses due to the mutual restriction b…

Response Generation

EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector

2024-11-04 · Deok-Hyeon Cho, Hyung-Seok Oh, Seung-bin Kim, Seong-Whan Lee

Emotional text-to-speech (TTS) technology has achieved significant progress in recent years; however, challenges remain owing to the inherent complexity of emotions and limitations of the available emotional speech datas…

DecoderEmotional Speech Synthesistext-to-speechText to Speech

Semi-supervised learning for continuous emotional intensity controllable speech synthesis with disentangled representations

2022-11-11 · Yoori Oh, Juheon Lee, Yoseob Han, Kyogu Lee

Recent text-to-speech models have reached the level of generating natural speech similar to what humans say. But there still have limitations in terms of expressiveness. The existing emotional speech synthesis models hav…

Emotional Speech SynthesisSpeech Synthesistext-to-speechText to Speech

UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech

2025-05-15 · Jiaxuan Liu, ZhenHua Ling

Recent neural codec language models have made great progress in the field of text-to-speech (TTS), but controllable emotional TTS still faces many challenges. Traditional methods rely on predefined discrete emotion label…

Emotional Speech SynthesisLanguage ModelingLanguage ModellingSpeech Synthesis+2