paper-with-me

홈 › Papers

Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought

2025-02-25 · Zhixian Zhao, Xinfa Zhu, Xinsheng Wang, Shuiyuan Wang, Xuelong Geng, Wenjie Tian, Lei Xie

Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable of comprehending diverse audio signal, performing audio analysis and generating textual responses. However, in speech emotion recognition (SER), ALMs often suffer from hallucinations, resulting in misclassifications or irrelevant outputs. To address these challenges, we propose C$^2$SER, a novel ALM designed to enhance the stability and accuracy of SER through Contextual perception and Chain of Thought (CoT). C$^2$SER integrates the Whisper encoder for semantic perception and Emotion2Vec-S for acoustic perception, where Emotion2Vec-S extends Emotion2Vec with semi-supervised learning to enhance emotional discrimination. Additionally, C$^2$SER employs a CoT approach, processing SER in a step-by-step manner while leveraging speech content and speaking styles to improve recognition. To further enhance stability, C$^2$SER introduces self-distillation from explicit CoT to implicit CoT, mitigating error accumulation and boosting recognition accuracy. Extensive experiments show that C$^2$SER outperforms existing popular ALMs, such as Qwen2-Audio and SECap, delivering more stable and precise emotion recognition. We release the training code, checkpoints, and test sets to facilitate further research.

📄 PDF Abstract BibTeX arXiv:2502.18186

Code (1)

zxzhao0/c2ser 공식 구현 pytorch

Tasks

Emotion RecognitionLanguage ModelingLanguage ModellingSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

2026-01-30 · Li Zhou, Hao Jiang, Junjie Li, Tianrui Wang 외 arxiv

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language…

Speech Synthesis

EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering

2025-08-05 · Tianxin Xie, Shan Yang, Chenxing Li, Dong Yu 외 arxiv

Text-to-speech (TTS) has shown great progress in recent years. However, most existing TTS systems offer only coarse and rigid emotion control, typically via discrete emotion labels or a carefully crafted and detailed emo…

Continuous Control

A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

2026-07-01 · Siyi Wang, James Bailey, Ting Dang arxiv

While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood. We present the first comparati…

Speech Synthesis

CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering

2026-02-03 · Siyi Wang, Shihong Tan, Siyi Liu, Hong Jia 외 arxiv

Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech sys…

Deep scattering network for speech emotion recognition

2021-05-11 · Premjeet Singh, Goutam Saha, Md Sahidullah

This paper introduces scattering transform for speech emotion recognition (SER). Scattering transform generates feature representations which remain stable to deformations and shifting in time and frequency without much …

Emotion RecognitionSpeech Emotion Recognition