paper-with-me

Papers

CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models

2026-03-30 · Kesheng Chen, Yamin Hu, Qi Zhou, Zhenqian Zhu, Wenjian Luo arxiv

Vision-language models (VLMs) achieve strong performance on many benchmarks, yet a basic reliability question remains underexplored: when visual evidence conflicts with commonsense, do models follow what is shown or what commonsense suggests? A characteristic failure in this setting is that the model overrides visual evidence and outputs the commonsense alternative. We term this phenomenon \textbf{commonsense-driven hallucination} (CDH). To evaluate it, we introduce \textbf{CDH-Bench}, a benchmark designed to create explicit \textbf{visual evidence--commonsense conflicts}. CDH-Bench covers three dimensions: \textit{counting anomalies}, \textit{relational anomalies}, and \textit{attribute anomalies}. We evaluate frontier VLMs under \textit{binary Question Answering (QA)} and \textit{multiple-choice QA}, and report metrics including \textit{Counterfactual Accuracy} (CF-Acc), \textit{Commonsense Accuracy} (CS-Acc), \textit{Counterfactual Accuracy Drop} (CFAD), \textit{Commonsense Collapse Rate} (CCR), and \textit{Relative Prior Dependency} (RPD). Results show that even strong models remain vulnerable to prior-driven normalization under visual evidence--commonsense conflict. CDH-Bench provides a controlled diagnostic of visual fidelity under visual evidence--commonsense conflict.

📄 PDF Abstract BibTeX arXiv:2603.27982

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding

2025-05-02 · Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin 외

Synthetic video generation has gained significant attention for its realism and broad applications, but remains prone to violations of common sense and physical laws. This highlights the need for reliable abnormality det…

Anomaly DetectionCommon Sense ReasoningHallucinationMVBench+2

CommonWhy: A Dataset for Evaluating Entity-Based Causal Commonsense Reasoning in Large Language Models

2026-05-13 · Armin Toroghi, Faeze Moradi Kalarde, Scott Sanner arxiv

To effectively interact with the real world, Large Language Models (LLMs) require entity-based commonsense reasoning, a challenging task that necessitates integrating factual knowledge about specific entities with common…

Graph Question Answering

CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge

2025-04-20 · Armin Toroghi, Willis Guo, Scott Sanner

The rise of Large Language Models (LLMs) has redefined the AI landscape, particularly due to their ability to encode factual and commonsense knowledge, and their outstanding performance in tasks requiring reasoning. Desp…

Claim VerificationGraph Question AnsweringQuestion Answering

LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning

2026-01-23 · Obed Junias, Maria Leonor Pacheco arxiv

Commonsense reasoning often involves evaluating multiple plausible interpretations rather than selecting a single atomic answer, yet most benchmarks rely on single-label evaluation, obscuring whether statements are joint…

ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation

2026-05-28 · Shizhe Zhou, Bohan Jia, Kai Wu, Yan Shen 외 arxiv

While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinations, producing responses that are inconsistent with the visual inpu…