Interpretable Multimodal Emotion Recognition using Hybrid Fusion of Speech and Image Data
This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been developed to identify the important speech & image features leading to the prediction of particular emotion classes. The proposed system's architecture has been determined through intensive ablation studies. It fuses the speech & image features and then combines speech, image, and intermediate fusion outputs. The proposed interpretability technique incorporates the divide & conquer approach to compute shapely values denoting each speech & image feature's importance. We have also constructed a large-scale dataset (IIT-R SIER dataset), consisting of speech utterances, corresponding images, and class labels, i.e., 'anger,' 'happy,' 'hate,' and 'sad.' The proposed system has achieved 83.29% accuracy for emotion recognition. The enhanced performance of the proposed system advocates the importance of utilizing complementary information from multiple modalities for emotion recognition.
Code (1)
Tasks
Emotion RecognitionMultimodal Emotion RecognitionSimilar Papers 제목 키워드 기반
VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
This paper proposes a multimodal emotion recognition system, VIsual Spoken Textual Additive Net (VISTANet), to classify emotions reflected by input containing image, speech, and text into discrete classes. A new interpre…
Emotion RecognitionMultimodal Emotion RecognitionMCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
Multimodal emotion recognition is crucial for future human-computer interaction. However, accurate emotion recognition still faces significant challenges due to differences between different modalities and the difficulty…
Multimodal Emotion RecognitionMultimodal Prompt Transformer with Hybrid Contrastive Learning for Emotion Recognition in Conversation
Emotion Recognition in Conversation (ERC) plays an important role in driving the development of human-machine interaction. Emotions can exist in multiple modalities, and multimodal ERC mainly faces two problems: (1) the …
Contrastive LearningEmotion RecognitionEmotion Recognition in ConversationFusion with Hierarchical Graphs for Mulitmodal Emotion Recognition
Automatic emotion recognition (AER) based on enriched multimodal inputs, including text, speech, and visual clues, is crucial in the development of emotionally intelligent machines. Although complex modality relationship…
Emotion ClassificationEmotion Recognitiongraph constructionMultimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph
Analyzing human multimodal language is an emerging area of research in NLP. Intrinsically this language is multimodal (heterogeneous), sequential and asynchronous; it consists of the language (words), visual (expressions…
Emotion RecognitionLanguage ModelingLanguage ModellingMultimodal Sentiment Analysis+2