paper-with-me

홈 › Papers

Leveraging Label Information for Multimodal Emotion Recognition

2023-09-05 · Peiying Wang, Sunlu Zeng, Junqing Chen, Lu Fan, Meng Chen, Youzheng Wu, Xiaodong He

Multimodal emotion recognition (MER) aims to detect the emotional status of a given expression by combining the speech and text information. Intuitively, label information should be capable of helping the model locate the salient tokens/frames relevant to the specific emotion, which finally facilitates the MER task. Inspired by this, we propose a novel approach for MER by leveraging label information. Specifically, we first obtain the representative label embeddings for both text and speech modalities, then learn the label-enhanced text/speech representations for each utterance via label-token and label-frame interactions. Finally, we devise a novel label-guided attentive fusion module to fuse the label-aware text and speech representations for emotion classification. Extensive experiments were conducted on the public IEMOCAP dataset, and experimental results demonstrate that our proposed approach outperforms existing baselines and achieves new state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2309.02106

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion ClassificationEmotion RecognitionMultimodal Emotion Recognition

Similar Papers 제목 키워드 기반

Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout

2024-09-11 · Anbin QI, Zhongliang Liu, Xinyong Zhou, Jinba Xiao 외

In this paper, we present our solution for the Second Multimodal Emotion Recognition Challenge Track 1(MER2024-SEMI). To enhance the accuracy and generalization performance of emotion recognition, we propose several meth…

Emotion RecognitionMultimodal Emotion RecognitionPrompt Learning

Leveraging Label Potential for Enhanced Multimodal Emotion Recognition

2025-04-07 · Xuechun Shao, Yinfeng Yu, Liejun Wang

Multimodal emotion recognition (MER) seeks to integrate various modalities to predict emotional states accurately. However, most current research focuses solely on the fusion of audio and text features, overlooking the v…

Emotion RecognitionMultimodal Emotion Recognition

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

2025-07-25 · Hsuan-Yu Wang, Pei-Ying Lee, Berlin Chen arxiv

In this paper, we investigate the impact of incorporating timestamp-based alignment between Automatic Speech Recognition (ASR) transcripts and Speaker Diarization (SD) outputs on Speech Emotion Recognition (SER) accuracy…

Multimodal Emotion RecognitionSpeech Emotion RecognitionSpeaker DiarizationSpeech Recognition

TiCAL:Typicality-Based Consistency-Aware Learning for Multimodal Emotion Recognition

2025-11-19 · Wen Yin, Siyu Zhan, Cencen Liu, Xin Hu 외 arxiv

Multimodal Emotion Recognition (MER) aims to accurately identify human emotional states by integrating heterogeneous modalities such as visual, auditory, and textual data. Existing approaches predominantly rely on unifie…

Multimodal Emotion Recognition

A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion

2025-02-12 · Wei Dai, Dequan Zheng, Feng Yu, Yanrong Zhang 외

With the advancement of artificial intelligence and computer vision technologies, multimodal emotion recognition has become a prominent research topic. However, existing methods face challenges such as heterogeneous data…

Contrastive LearningEmotion RecognitionMultimodal Emotion Recognition