paper-with-me

홈 › Papers

Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models

2026-03-18 · Xiutian Zhao, Ismail Rasim Ulgen, Philipp Koehn, Björn Schuller, Berrak Sisman arxiv

Large audio-language models (LALMs) can produce expressive speech, yet reliable emotion control remains elusive: conversions often miss the target affect and may degrade linguistic fidelity through refusals, hallucinations, or paraphrase. We present, to our knowledge, the first neuron-level study of emotion control in speech-generative LALMs and demonstrate that compact emotion-sensitive neurons (ESNs) are causally actionable, enabling training-free emotion steering at inference time. ESNs are identified via success-filtered activation aggregation enforcing both emotion realization and content preservation. Across three LALMs (Qwen2.5-Omni-7B, MiniCPM-o 4.5, Kimi-Audio), ESN interventions yield emotion-specific gains that generalize to unseen speakers and are supported by automatic and human evaluation. Controllability depends on selector design, mask sparsity, filtering, and intervention strength. Our results establish a mechanistic framework for training-free emotion control in speech generation.

📄 PDF Abstract BibTeX arXiv:2603.17231

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations

2024-12-09 · Weizhen Bian, Yubo Zhou, Kaitai Zhang, Xiaohan Gu

Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human …

text-to-speechText to Speech

SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion Diarization

2026-04-14 · Farzaneh Jafari, Stefano Berretti, Anup Basu arxiv

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely…

Talking Head GenerationEmotion Recognition

QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis

2023-03-14 · Haobin Tang, xulong Zhang, Jianzong Wang, Ning Cheng 외

Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and cont…

Emotional Speech SynthesisSentenceSpeech Synthesistext-to-speech+1

Semi-supervised learning for continuous emotional intensity controllable speech synthesis with disentangled representations

2022-11-11 · Yoori Oh, Juheon Lee, Yoseob Han, Kyogu Lee

Recent text-to-speech models have reached the level of generating natural speech similar to what humans say. But there still have limitations in terms of expressiveness. The existing emotional speech synthesis models hav…

Emotional Speech SynthesisSpeech Synthesistext-to-speechText to Speech

Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models

2026-01-06 · Xiutian Zhao, Björn Schuller, Berrak Sisman arxiv

Emotion is a central dimension of spoken communication, yet, we still lack a mechanistic account of how modern large audio-language models (LALMs) encode it internally. We present the first neuron-level interpretability …

Emotion Recognition