paper-with-me

홈 › Papers

Interpretable Multimodal Emotion Recognition using Hybrid Fusion of Speech and Image Data

2022-08-25 · Puneet Kumar, Sarthak Malik, Balasubramanian Raman

This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been developed to identify the important speech & image features leading to the prediction of particular emotion classes. The proposed system's architecture has been determined through intensive ablation studies. It fuses the speech & image features and then combines speech, image, and intermediate fusion outputs. The proposed interpretability technique incorporates the divide & conquer approach to compute shapely values denoting each speech & image feature's importance. We have also constructed a large-scale dataset (IIT-R SIER dataset), consisting of speech utterances, corresponding images, and class labels, i.e., 'anger,' 'happy,' 'hate,' and 'sad.' The proposed system has achieved 83.29% accuracy for emotion recognition. The enhanced performance of the proposed system advocates the importance of utilizing complementary information from multiple modalities for emotion recognition.

📄 PDF Abstract BibTeX arXiv:2208.11868

Code (1)

mintelligence-group/speechimg_emorec 공식 구현

Tasks

Emotion RecognitionMultimodal Emotion Recognition

Similar Papers 제목 키워드 기반

VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition

2022-08-24 · Puneet Kumar, Sarthak Malik, Balasubramanian Raman, Xiaobai Li

This paper proposes a multimodal emotion recognition system, VIsual Spoken Textual Additive Net (VISTANet), to classify emotions reflected by input containing image, speech, and text into discrete classes. A new interpre…

Emotion RecognitionMultimodal Emotion Recognition

MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition

2025-10-28 · Haoyang Zhang, Zhou Yang, Ke Sun, Yucai Pang 외 arxiv

Multimodal emotion recognition is crucial for future human-computer interaction. However, accurate emotion recognition still faces significant challenges due to differences between different modalities and the difficulty…

Multimodal Emotion Recognition

Multimodal Prompt Transformer with Hybrid Contrastive Learning for Emotion Recognition in Conversation

2023-10-04 · Shihao Zou, Xianying Huang, Xudong Shen

Emotion Recognition in Conversation (ERC) plays an important role in driving the development of human-machine interaction. Emotions can exist in multiple modalities, and multimodal ERC mainly faces two problems: (1) the …

Contrastive LearningEmotion RecognitionEmotion Recognition in Conversation

Fusion with Hierarchical Graphs for Mulitmodal Emotion Recognition

2021-09-15 · Shuyun Tang, Zhaojie Luo, Guoshun Nan, Yuichiro Yoshikawa 외

Automatic emotion recognition (AER) based on enriched multimodal inputs, including text, speech, and visual clues, is crucial in the development of emotionally intelligent machines. Although complex modality relationship…

Emotion ClassificationEmotion Recognitiongraph construction

Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

2018-07-01 · ACL 2018 7 · AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria 외

Analyzing human multimodal language is an emerging area of research in NLP. Intrinsically this language is multimodal (heterogeneous), sequential and asynchronous; it consists of the language (words), visual (expressions…

Emotion RecognitionLanguage ModelingLanguage ModellingMultimodal Sentiment Analysis+2