Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches often treat emotion understanding as a classification problem, offering little insight into the underlying rationale behind predictions. In this work, we explore emotion reasoning, a strategy that leverages the generative capabilities of AudioLLMs to enhance emotion recognition by producing semantically aligned, evidence-grounded explanations. To support this in multitask AudioLLMs, we introduce a unified framework combining reasoning-augmented data supervision, dual-encoder architecture, and task-alternating training. This approach enables AudioLLMs to effectively learn different tasks while incorporating emotional reasoning. Experiments on IEMOCAP and MELD show that our approach not only improves emotion prediction accuracy but also enhances the coherence and evidential grounding of the generated responses.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion Recognitionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Color-based Emotion Representation for Speech Emotion Recognition
Speech emotion recognition (SER) has traditionally relied on categorical or dimensional labels. However, this technique is limited in representing both the diversity and interpretability of emotions. To overcome this lim…
Speech Emotion RecognitionEmotion ClassificationJointly Predicting Emotion, Age, and Country Using Pre-Trained Acoustic Embedding
In this paper, we demonstrated the benefit of using pre-trained model to extract acoustic embedding to jointly predict (multitask learning) three tasks: emotion, age, and native country. The pre-trained model was trained…
regressionLearning Representations of Emotional Speech with Deep Convolutional Generative Adversarial Networks
Automatically assessing emotional valence in human speech has historically been a difficult task for machine learning algorithms. The subtle changes in the voice of the speaker that are indicative of positive or negative…
BIG-bench Machine LearningGeneral ClassificationGenerative Adversarial NetworkRepresentation LearningMultitask Learning for Emotionally Analyzing Sexual Abuse Disclosures
The {\#}MeToo movement on social media platforms initiated discussions over several facets of sexual harassment in our society. Prior work by the NLP community for automated identification of the narratives related to se…
ClassificationEmotion ClassificationHate Speech DetectionTransfer LearningLearning Spontaneity to Improve Emotion Recognition In Speech
We investigate the effect and usefulness of spontaneity (i.e. whether a given speech is spontaneous or not) in speech in the context of emotion recognition. We hypothesize that emotional content in speech is interrelated…
Emotion RecognitionSpeech Emotion Recognition