paper-with-me

Papers

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness

2026-05-01 · Yueru Sun, Yimeng Zhang, Haoyu Gu, Nuo Chen, Dong She, Xianrong Yao, Yang Gao, Zhanpeng Jin arxiv

Multimodal Emotion Recognition (MER) is critical for interpreting real-world interactions. While Multimodal Large Language Models (MLLM) have shown promise in MER, their internal decision-making mechanisms under modality conflict and missingness remain largely underexplored. In this paper, to systematically investigate these behaviors, we introduce EmoMM, a comprehensive benchmark featuring modality-aligned, conflict, and missing subsets. Through extensive evaluation, we uncover a Video Contribution Collapse (VCC) phenomenon, where MLLM marginalize video evidence due to high token redundancy and modality preferences. To address this, we propose Conflict-aware Head-level Attention Steering (CHASE), a lightweight mechanism that detects modality conflicts and performs inference-time attention steering, effectively mitigating decision bias without retraining the backbone. Experimental results demonstrate that CHASE consistently improves performance across various settings, significantly enhancing the reliability of MLLM in complex affective scenarios.

📄 PDF Abstract BibTeX arXiv:2605.01024

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Emotion Recognition

Similar Papers 제목 키워드 기반

EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models

2025-02-06 · He Hu, Yucheng Zhou, Lianzhong You, Hongbo Xu 외

With the integration of Multimodal large language models (MLLMs) into robotic systems and various AI applications, embedding emotional intelligence (EI) capabilities into these models is essential for enabling robots to …

BenchmarkingEmotional IntelligenceEmotion Recognition

StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models

2024-08-31 · Yuxiang Guo, Faizan Siddiqui, Yang Zhao, Rama Chellappa 외

Predicting and reasoning how a video would make a human feel is crucial for developing socially intelligent systems. Although Multimodal Large Language Models (MLLMs) have shown impressive video understanding capabilitie…

Video Understanding

Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning

2025-08-02 · Zhiyuan Han, Beier Zhu, Yanlong Xu, Peipei Song 외 arxiv

Despite their strong performance in multimodal emotion reasoning, existing Multimodal Large Language Models (MLLMs) often overlook the scenarios involving emotion conflicts, where emotional cues from different modalities…

MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms

2024-02-21 · Yiqiao Jin, MinJe Choi, Gaurav Verma, Jindong Wang 외

Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emotions associated with interactions in onl…

BenchmarkingHate Speech DetectionMisinformation

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

2026-08-04 · Xiaolin Chen, Xuemeng Song, Wenhao Shi, Xianjing Han 외 arxiv

Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotion…

Explanation Generation