paper-with-me

홈 › Papers

ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Models

2024-09-14 · CVPR 2025 1 · Yahan Tu, Rui Hu, Jitao Sang

Hallucination poses a persistent challenge for multimodal large language models (MLLMs). However, existing benchmarks for evaluating hallucinations are generally static, which may overlook the potential risk of data contamination. To address this issue, we propose ODE, an open-set, dynamic protocol designed to evaluate object hallucinations in MLLMs at both the existence and attribute levels. ODE employs a graph-based structure to represent real-world object concepts, their attributes, and the distributional associations between them. This structure facilitates the extraction of concept combinations based on diverse distributional criteria, generating varied samples for structured queries that evaluate hallucinations in both generative and discriminative tasks. Through the generation of new samples, dynamic concept combinations, and varied distribution frequencies, ODE mitigates the risk of data contamination and broadens the scope of evaluation. This protocol is applicable to both general and specialized scenarios, including those with limited data. Experimental results demonstrate the effectiveness of our protocol, revealing that MLLMs exhibit higher hallucination rates when evaluated with ODE-generated samples, which indicates potential data contamination. Furthermore, these generated samples aid in analyzing hallucination patterns and fine-tuning models, offering an effective approach to mitigating hallucinations in MLLMs.

📄 PDF Abstract BibTeX arXiv:2409.09318

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeHallucination

Similar Papers 제목 키워드 기반

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

2025-08-18 · Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to…

Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models

2024-08-18 · Kening Zheng, Junkai Chen, Yibo Yan, Xin Zou 외

Hallucination issues continue to affect multimodal large language models (MLLMs), with existing research mainly addressing object-level or attribute-level hallucinations, neglecting the more complex relation hallucinatio…

AttributeHallucinationHallucination EvaluationRelation

EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models

2025-05-16 · Bohao Xing, Xin Liu, Guoying Zhao, Chengyu Liu 외

Emotion understanding is a critical yet challenging task. Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced their capabilities in this area. However, MLLMs often suffer from hallucin…

Hallucination

Unified Hallucination Detection for Multimodal Large Language Models

2024-02-05 · Xiang Chen, Chenxi Wang, Yida Xue, Ningyu Zhang 외

Despite significant strides in multimodal tasks, Multimodal Large Language Models (MLLMs) are plagued by the critical issue of hallucination. The reliable detection of such hallucinations in MLLMs has, therefore, become …

Hallucination

VIGIL: Tackling Hallucination Detection in Image Recontextualization

2026-02-16 · Joanna Wojciechowicz, Maria Łubniewska, Jakub Antczak, Justyna Baczyńska 외 arxiv

We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), the first benchmark dataset and framework providing a fine-grained categorization of hallucinations in the multimodal image recontextualization …