MMER: Multimodal Multi-task Learning for Speech Emotion Recognition
In this paper, we propose MMER, a novel Multimodal Multi-task learning approach for Speech Emotion Recognition. MMER leverages a novel multimodal network based on early-fusion and cross-modal self-attention between text and acoustic modalities and solves three novel auxiliary tasks for learning emotion recognition from spoken utterances. In practice, MMER outperforms all our baselines and achieves state-of-the-art performance on the IEMOCAP benchmark. Additionally, we conduct extensive ablation studies and results analysis to prove the effectiveness of our proposed approach.
Code (1)
Tasks
Emotion RecognitionMultimodal Emotion RecognitionMulti-Task LearningSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
AIVA: An AI-based Virtual Companion for Emotion-aware Interaction
Recent advances in Large Language Models (LLMs) have significantly improved natural language understanding and generation, enhancing Human-Computer Interaction (HCI). However, LLMs are limited to unimodal text processing…
Natural Language UnderstandingContrastive LearningPrompt EngineeringLearning Alignment for Multimodal Emotion Recognition from Speech
Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or…
Emotion RecognitionMultimodal Emotion RecognitionSpeech Emotion Recognitionspeech-recognition+1Emotion Impacts Speech Recognition Performance
It has been established that the performance of speech recognition systems depends on multiple factors including the lexical content, speaker identity and dialect. Here we use three English datasets of acted emotion to d…
speech-recognitionSpeech RecognitionWavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
Speech emotion recognition (SER) remains a challenging yet crucial task due to the inherent complexity and diversity of human emotions. To address this problem, researchers attempt to fuse information from other modaliti…
DiversityEmotion RecognitionRepresentation LearningSpeech Emotion RecognitionJointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
Multimodal emotion recognition from speech is an important area in affective computing. Fusing multiple data modalities and learning representations with limited amounts of labeled data is a challenging task. In this pap…