paper-with-me

홈 › Papers

Revisiting Disentanglement and Fusion on Modality and Context in Conversational Multimodal Emotion Recognition

2023-08-08 · Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li

It has been a hot research topic to enable machines to understand human emotions in multimodal contexts under dialogue scenarios, which is tasked with multimodal emotion analysis in conversation (MM-ERC). MM-ERC has received consistent attention in recent years, where a diverse range of methods has been proposed for securing better task performance. Most existing works treat MM-ERC as a standard multimodal classification problem and perform multimodal feature disentanglement and fusion for maximizing feature utility. Yet after revisiting the characteristic of MM-ERC, we argue that both the feature multimodality and conversational contextualization should be properly modeled simultaneously during the feature disentanglement and fusion steps. In this work, we target further pushing the task performance by taking full consideration of the above insights. On the one hand, during feature disentanglement, based on the contrastive learning technique, we devise a Dual-level Disentanglement Mechanism (DDM) to decouple the features into both the modality space and utterance space. On the other hand, during the feature fusion stage, we propose a Contribution-aware Fusion Mechanism (CFM) and a Context Refusion Mechanism (CRM) for multimodal and context integration, respectively. They together schedule the proper integrations of multimodal and context features. Specifically, CFM explicitly manages the multimodal feature contributions dynamically, while CRM flexibly coordinates the introduction of dialogue contexts. On two public MM-ERC datasets, our system achieves new state-of-the-art performance consistently. Further analyses demonstrate that all our proposed mechanisms greatly facilitate the MM-ERC task by making full use of the multimodal and context features adaptively. Note that our proposed methods have the great potential to facilitate a broader range of other conversational multimodal tasks.

📄 PDF Abstract BibTeX arXiv:2308.04502

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningDisentanglementEmotion RecognitionEmotion Recognition in ConversationMultimodal Emotion Recognition

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Revisiting Multi-modal Emotion Learning with Broad State Space Models and Probability-guidance Fusion

2024-04-27 · Yuntao Shou, Tao Meng, FuChen Zhang, Nan Yin 외

Multi-modal Emotion Recognition in Conversation (MERC) has received considerable attention in various fields, e.g., human-computer interaction and recommendation systems. Most existing works perform feature disentangleme…

DisentanglementEmotion ClassificationEmotion RecognitionEmotion Recognition in Conversation+3

Beyond Whole Dialogue Modeling: Contextual Disentanglement for Conversational Recommendation

2025-04-24 · Guojia An, Jie Zou, Jiwei Wei, Chaoning Zhang 외

Conversational recommender systems aim to provide personalized recommendations by analyzing and utilizing contextual information related to dialogue. However, existing methods typically model the dialogue context as a wh…

Conversational RecommendationcounterfactualCounterfactual InferenceDisentanglement+3

Revisiting Conversation Discourse for Dialogue Disentanglement

2023-06-06 · Bobo Li, Hao Fei, Fei Li, Shengqiong Wu 외

Dialogue disentanglement aims to detach the chronologically ordered utterances into several independent sessions. Conversation utterances are essentially organized and described by the underlying discourse, and thus dial…

AttributeDisentanglement

Revisiting Cross-Modal Knowledge Distillation: A Disentanglement Approach for RGBD Semantic Segmentation

2025-05-30 · Roger Ferrod, Cássio F. Dantas, Luigi di Caro, Dino Ienco

Multi-modal RGB and Depth (RGBD) data are predominant in many domains such as robotics, autonomous driving and remote sensing. The combination of these multi-modal data enhances environmental perception by providing 3D s…

Autonomous DrivingContrastive LearningData AugmentationDisentanglement+3

Disentangled Dual-Branch Graph Learning for Conversational Emotion Recognition

2026-04-03 · Chengling Guo, Yuntao Shou, Tao Meng, Wei Ai 외 arxiv

Multimodal emotion recognition in conversations aims to infer utterance-level emotions by jointly modeling textual, acoustic, and visual cues within context. Despite recent progress, key challenges remain, including redu…

Multimodal Emotion RecognitionGraph Neural NetworkGraph Learning