paper-with-me

Papers

A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

2026-07-01 · Siyi Wang, James Bailey, Ting Dang arxiv

While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood. We present the first comparative study of speech language model (SLM) and conditional flow-matching (CFM) modules as activation steering sites for mixed emotion speech synthesis. We first characterize emotion representations using linear probing and local intrinsic dimensionality (LID), and then evaluate single-site and joint steering for mixed-emotion synthesis. Our results show that SLM offers a clean, low-dimensional emotion-specific subspace with strong speaker--emotion disentanglement, while CFM exhibitspoor cross-speaker generalization due to speaker--emotion entanglement. Joint steering increases emotion intensity but degrades proportional control and speech quality on in-distribution data. These findings provide practical guidance for multi-site activation steering in hybrid TTS systems and highlight the importance of representation geometry in controllable speech generation.

📄 PDF Abstract BibTeX arXiv:2607.00946

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

2026-02-03 · Siyi Wang, Shihong Tan, Siyi Liu, Hong Jia 외 arxiv

Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech sys…

EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering

2025-08-05 · Tianxin Xie, Shan Yang, Chenxing Li, Dong Yu 외 arxiv

Text-to-speech (TTS) has shown great progress in recent years. However, most existing TTS systems offer only coarse and rigid emotion control, typically via discrete emotion labels or a carefully crafted and detailed emo…

Continuous Control

Why Are Some Emotions Harder for LLMs? Uncovering the Causal Mechanisms of Emotion Inference via Sparse Autoencoders

2026-04-28 · Bangzhao Shu, Arinjay Singh, Mai ElSherief arxiv

Large language models (LLMs) are increasingly used in emotionally sensitive human-AI applications, where reliable emotion detection is essential. However, their emotion recognition abilities remain uneven: models often p…

Emotion Recognition

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

2026-01-30 · Li Zhou, Hao Jiang, Junjie Li, Tianrui Wang 외 arxiv

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language…

Speech Synthesis

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

2026-08-06 · Mehrshad Saadatinia, Parsa Razmara, Ardalan Aryashad, Ali Abbasi 외 arxiv

Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single-layer inter…