paper-with-me

홈 › Papers

Interpretability for Multimodal Emotion Recognition using Concept Activation Vectors

2022-02-02 · Ashish Ramayee Asokan, Nidarshan Kumar, Anirudh Venkata Ragam, Shylaja S Sharath

Multimodal Emotion Recognition refers to the classification of input video sequences into emotion labels based on multiple input modalities (usually video, audio and text). In recent years, Deep Neural networks have shown remarkable performance in recognizing human emotions, and are on par with human-level performance on this task. Despite the recent advancements in this field, emotion recognition systems are yet to be accepted for real world setups due to the obscure nature of their reasoning and decision-making process. Most of the research in this field deals with novel architectures to improve the performance for this task, with a few attempts at providing explanations for these models' decisions. In this paper, we address the issue of interpretability for neural networks in the context of emotion recognition using Concept Activation Vectors (CAVs). To analyse the model's latent space, we define human-understandable concepts specific to Emotion AI and map them to the widely-used IEMOCAP multimodal database. We then evaluate the influence of our proposed concepts at multiple layers of the Bi-directional Contextual LSTM (BC-LSTM) network to show that the reasoning process of neural networks for emotion recognition can be represented using human-understandable concepts. Finally, we perform hypothesis testing on our proposed concepts to show that they are significant for interpretability of this task.

📄 PDF Abstract BibTeX arXiv:2202.01072

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingEmotion RecognitionMultimodal Emotion Recognition

Methods 이 논문이 사용한 방법론

Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

GatedxLSTM: A Multimodal Affective Computing Approach for Emotion Recognition in Conversations

2025-03-26 · Yupei Li, Qiyang Sun, Sunil Munthumoduku Krishna Murthy, Emran Alturki 외

Affective Computing (AC) is essential for advancing Artificial General Intelligence (AGI), with emotion recognition serving as a key component. However, human emotions are inherently dynamic, influenced not only by an in…

cross-modal alignmentEmotion ClassificationEmotion RecognitionEmotion Recognition in Conversation+1

Interpretable Multimodal Emotion Recognition using Hybrid Fusion of Speech and Image Data

2022-08-25 · Puneet Kumar, Sarthak Malik, Balasubramanian Raman

This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability techniq…

Emotion RecognitionMultimodal Emotion Recognition

Interpretable Multimodal Emotion Recognition using Facial Features and Physiological Signals

2023-06-05 · Puneet Kumar, Xiaobai Li

This paper aims to demonstrate the importance and feasibility of fusing multimodal information for emotion recognition. It introduces a multimodal framework for emotion understanding by fusing the information from visual…

Emotion ClassificationEmotion RecognitionFeature ImportanceMultimodal Emotion Recognition

Do Vision-Language Pretrained Models Learn Composable Primitive Concepts?

2022-03-31 · Tian Yun, Usha Bhalla, Ellie Pavlick, Chen Sun

Vision-language (VL) pretrained models have achieved impressive performance on multimodal reasoning and zero-shot recognition tasks. Many of these VL models are pretrained on unlabeled image and caption pairs from the in…

Fine-Grained Visual RecognitionMultimodal ReasoningZero-Shot Learning

VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition

2022-08-24 · Puneet Kumar, Sarthak Malik, Balasubramanian Raman, Xiaobai Li

This paper proposes a multimodal emotion recognition system, VIsual Spoken Textual Additive Net (VISTANet), to classify emotions reflected by input containing image, speech, and text into discrete classes. A new interpre…

Emotion RecognitionMultimodal Emotion Recognition