GatedxLSTM: A Multimodal Affective Computing Approach for Emotion Recognition in Conversations
Affective Computing (AC) is essential for advancing Artificial General Intelligence (AGI), with emotion recognition serving as a key component. However, human emotions are inherently dynamic, influenced not only by an individual's expressions but also by interactions with others, and single-modality approaches often fail to capture their full dynamics. Multimodal Emotion Recognition (MER) leverages multiple signals but traditionally relies on utterance-level analysis, overlooking the dynamic nature of emotions in conversations. Emotion Recognition in Conversation (ERC) addresses this limitation, yet existing methods struggle to align multimodal features and explain why emotions evolve within dialogues. To bridge this gap, we propose GatedxLSTM, a novel speech-text multimodal ERC model that explicitly considers voice and transcripts of both the speaker and their conversational partner(s) to identify the most influential sentences driving emotional shifts. By integrating Contrastive Language-Audio Pretraining (CLAP) for improved cross-modal alignment and employing a gating mechanism to emphasise emotionally impactful utterances, GatedxLSTM enhances both interpretability and performance. Additionally, the Dialogical Emotion Decoder (DED) refines emotion predictions by modelling contextual dependencies. Experiments on the IEMOCAP dataset demonstrate that GatedxLSTM achieves state-of-the-art (SOTA) performance among open-source methods in four-class emotion classification. These results validate its effectiveness for ERC applications and provide an interpretability analysis from a psychological perspective.
Code (0)
등록된 구현이 없습니다.
Tasks
cross-modal alignmentEmotion ClassificationEmotion RecognitionEmotion Recognition in ConversationMultimodal Emotion RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Recent Trends of Multimodal Affective Computing: A Survey from NLP Perspective
Multimodal affective computing (MAC) has garnered increasing attention due to its broad applications in analyzing human behaviors and intentions, especially in text-dominated multimodal affective computing field. This su…
Aspect-Based Sentiment AnalysisEmotion RecognitionEmotion Recognition in ConversationMultimodal Emotion Recognition+3Affective Video Content Analysis: Decade Review and New Perspectives
Video content is rich in semantics and has the ability to evoke various emotions in viewers. In recent years, with the rapid development of affective computing and the explosive growth of visual data, affective video con…
Emotional IntelligenceEmotion RecognitionFacial Expression RecognitionVideo Emotion RecognitionModeling emotion in complex stories: the Stanford Emotional Narratives Dataset
Human emotions unfold over time, and more affective computing research has to prioritize capturing this crucial component of real-world affect. Modeling dynamic emotional stimuli requires solving the twin challenges of t…
Emotion RecognitionTime SeriesTime Series AnalysisEmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding
Sentiment and emotion understanding are essential to applications such as human-computer interaction and depression detection. While Multimodal Large Language Models (MLLMs) demonstrate robust general capabilities, they …
Depression DetectionEmotion-Cause Pair ExtractionEmotion RecognitionFacial Expression Recognition+3Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition
This article presents our results for the eighth Affective Behavior Analysis in-the-wild (ABAW) competition.Multimodal emotion recognition (ER) has important applications in affective computing and human-computer interac…
Emotion RecognitionMultimodal Emotion Recognition