paper-with-me

홈 › Papers

Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs

2025-06-07 · Wenyu Zhang, Yingxu He, Geyu Lin, Zhuohan Liu, Shuo Sun, Bin Wang, Xunlong Zou, Jeremy H. M. Wong, Qiongqiong Wang, Hardik B. Sailor, Nancy F. Chen, Ai Ti Aw

Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches often treat emotion understanding as a classification problem, offering little insight into the underlying rationale behind predictions. In this work, we explore emotion reasoning, a strategy that leverages the generative capabilities of AudioLLMs to enhance emotion recognition by producing semantically aligned, evidence-grounded explanations. To support this in multitask AudioLLMs, we introduce a unified framework combining reasoning-augmented data supervision, dual-encoder architecture, and task-alternating training. This approach enables AudioLLMs to effectively learn different tasks while incorporating emotional reasoning. Experiments on IEMOCAP and MELD show that our approach not only improves emotion prediction accuracy but also enhances the coherence and evidential grounding of the generated responses.

📄 PDF Abstract BibTeX arXiv:2506.06820

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Color-based Emotion Representation for Speech Emotion Recognition

2026-02-18 · Ryotaro Nagase, Ryoichi Takashima, Yoichi Yamashita arxiv

Speech emotion recognition (SER) has traditionally relied on categorical or dimensional labels. However, this technique is limited in representing both the diversity and interpretability of emotions. To overcome this lim…

Speech Emotion RecognitionEmotion Classification

Jointly Predicting Emotion, Age, and Country Using Pre-Trained Acoustic Embedding

2022-07-21 · Bagus Tris Atmaja, Zanjabila, Akira Sasou

In this paper, we demonstrated the benefit of using pre-trained model to extract acoustic embedding to jointly predict (multitask learning) three tasks: emotion, age, and native country. The pre-trained model was trained…

regression

Learning Representations of Emotional Speech with Deep Convolutional Generative Adversarial Networks

2017-04-22 · Jonathan Chang, Stefan Scherer

Automatically assessing emotional valence in human speech has historically been a difficult task for machine learning algorithms. The subtle changes in the voice of the speaker that are indicative of positive or negative…

BIG-bench Machine LearningGeneral ClassificationGenerative Adversarial NetworkRepresentation Learning

Multitask Learning for Emotionally Analyzing Sexual Abuse Disclosures

2021-06-01 · NAACL 2021 4 · Ramit Sawhney, Puneet Mathur, Taru Jain, Akash Kumar Gautam 외

The {\#}MeToo movement on social media platforms initiated discussions over several facets of sexual harassment in our society. Prior work by the NLP community for automated identification of the narratives related to se…

ClassificationEmotion ClassificationHate Speech DetectionTransfer Learning

Learning Spontaneity to Improve Emotion Recognition In Speech

2017-12-12 · Karttikeya Mangalam, Tanaya Guha

We investigate the effect and usefulness of spontaneity (i.e. whether a given speech is spontaneous or not) in speech in the context of emotion recognition. We hypothesize that emotional content in speech is interrelated…

Emotion RecognitionSpeech Emotion Recognition