paper-with-me

홈 › Papers

HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning

2024-07-22 · Zhecan Wang, Garrett Bingham, Adams Yu, Quoc Le, Thang Luong, Golnaz Ghiasi

Hallucination has been a major problem for large language models and remains a critical challenge when it comes to multimodality in which vision-language models (VLMs) have to deal with not just textual but also visual inputs. Despite rapid progress in VLMs, resources for evaluating and addressing multimodal hallucination are limited and mostly focused on evaluation. This work introduces HaloQuest, a novel visual question answering dataset that captures various aspects of multimodal hallucination such as false premises, insufficient contexts, and visual challenges. A novel idea from HaloQuest is to leverage synthetic images, apart from real ones, to enable dataset creation at scale. With over 7.7K examples spanning across a wide variety of categories, HaloQuest was designed to be both a challenging benchmark for VLMs and a fine-tuning dataset for advancing multimodal reasoning. Our experiments reveal that current models struggle with HaloQuest, with all open-source VLMs achieving below 36% accuracy. On the other hand, fine-tuning on HaloQuest significantly reduces hallucination rates while preserving performance on standard reasoning tasks. Our results discover that benchmarking with generated images is highly correlated (r=0.97) with real images. Last but not least, we propose a novel Auto-Eval mechanism that is highly correlated with human raters (r=0.99) for evaluating VLMs. In sum, this work makes concrete strides towards understanding, evaluating, and mitigating hallucination in VLMs, serving as an important step towards more reliable multimodal AI systems in the future.

📄 PDF Abstract BibTeX arXiv:2407.15680

Code (1)

google/haloquest 공식 구현

Tasks

BenchmarkingHallucinationMultimodal ReasoningQuestion AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model

2025-05-29 · Nokimul Hasan Arif, Shadman Rabby, Md Hefzul Hossain Papon, Sabbir Ahmed

Visual hallucinations in Large Language Models (LLMs), where the model generates responses that are inconsistent with the visual input, pose a significant challenge to their reliability, particularly in contexts where pr…

HallucinationLanguage ModelingLanguage ModellingMultimodal Reasoning+1

Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling

2025-04-17 · Tsung-Han Wu, HeeKyung Lee, Jiaxin Ge, Joseph E. Gonzalez 외

Vision-Language Models (VLMs) excel at visual understanding but often suffer from visual hallucinations, where they generate descriptions of nonexistent objects, actions, or concepts, posing significant risks in safety-c…

Hallucination

Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning

2024-12-15 · Shengqiong Wu, Hao Fei, Liangming Pan, William Yang Wang 외

Recent advancements in multimodal large language models (MLLMs) have shown unprecedented capabilities in advancing various vision-language tasks. However, MLLMs face significant challenges with hallucinations, and mislea…

Hallucination

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment

2026-02-28 · Yantao Li, Qiang Hui, Chenyang Yan, Kanzhi Cheng 외 arxiv

Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucin…

Reinforcement LearningMultimodal ReasoningVisual Reasoning

TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models

2023-08-07 · Wenqi Shao, Meng Lei, Yutao Hu, Peng Gao 외

Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated significant progress in tackling complex multimodal tasks. Among these cutting-edge developments, Google's Bard stands out for its remarkable …

HallucinationObject HallucinationVisual Reasoning