paper-with-me

Papers

EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models

2025-05-16 · Bohao Xing, Xin Liu, Guoying Zhao, Chengyu Liu, Xiaolan Fu, Heikki Kälviäinen

Emotion understanding is a critical yet challenging task. Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced their capabilities in this area. However, MLLMs often suffer from hallucinations, generating irrelevant or nonsensical content. To the best of our knowledge, despite the importance of this issue, there has been no dedicated effort to evaluate emotion-related hallucinations in MLLMs. In this work, we introduce EmotionHallucer, the first benchmark for detecting and analyzing emotion hallucinations in MLLMs. Unlike humans, whose emotion understanding stems from the interplay of biology and social learning, MLLMs rely solely on data-driven learning and lack innate emotional instincts. Fortunately, emotion psychology provides a solid foundation of knowledge about human emotions. Building on this, we assess emotion hallucinations from two dimensions: emotion psychology knowledge and real-world multimodal perception. To support robust evaluation, we utilize an adversarial binary question-answer (QA) framework, which employs carefully crafted basic and hallucinated pairs to assess the emotion hallucination tendencies of MLLMs. By evaluating 38 LLMs and MLLMs on EmotionHallucer, we reveal that: i) most current models exhibit substantial issues with emotion hallucinations; ii) closed-source models outperform open-source ones in detecting emotion hallucinations, and reasoning capability provides additional advantages; iii) existing models perform better in emotion psychology knowledge than in multimodal emotion perception. As a byproduct, these findings inspire us to propose the PEP-MEK framework, which yields an average improvement of 9.90% in emotion hallucination detection across selected models. Resources will be available at https://github.com/xxtars/EmotionHallucer.

📄 PDF Abstract BibTeX arXiv:2505.11405

Code (1)

xxtars/emotionhallucer 공식 구현

Tasks

Hallucination

Similar Papers 제목 키워드 기반

AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

2026-02-04 · Ashutosh Chaubey, Jiacheng Pang, Maksim Siniukov, Mohammad Soleymani arxiv

Emotion understanding is essential for building socially intelligent agents. Although recent multimodal large language models have shown strong performance on this task, two key challenges remain - spurious associations …

Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions

2025-09-19 · Hansol Park, Hoseong Ahn, Junwon Moon, Yejin Lee 외 arxiv

Hallucinations in multimodal models have been extensively studied using benchmarks that probe reliability in image-text query settings. However, the effect of spoken queries on multimodal hallucinations remains largely u…

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

2024-12-04 · CVPR 2025 1 · Chaoyu Li, Eun Woo Im, Pooyan Fazli

Multimodal large language models (MLLMs) have recently shown significant advancements in video understanding, excelling in content reasoning and instruction-following tasks. However, hallucination, where models generate …

HallucinationInstruction FollowingSemantic SimilaritySemantic Textual Similarity+1

The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

2024-10-16 · Sicong Leng, Yun Xing, Zesen Cheng, Yang Zhou 외

Recent advancements in large multimodal models (LMMs) have significantly enhanced performance across diverse tasks, with ongoing efforts to further integrate additional modalities such as video and audio. However, most e…

Hallucination

EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models

2025-02-06 · He Hu, Yucheng Zhou, Lianzhong You, Hongbo Xu 외

With the integration of Multimodal large language models (MLLMs) into robotic systems and various AI applications, embedding emotional intelligence (EI) capabilities into these models is essential for enabling robots to …

BenchmarkingEmotional IntelligenceEmotion Recognition