paper-with-me

홈 › Papers

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

2023-10-23 · CVPR 2024 1 · Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, Tianyi Zhou

We introduce HallusionBench, a comprehensive benchmark designed for the evaluation of image-context reasoning. This benchmark presents significant challenges to advanced large visual-language models (LVLMs), such as GPT-4V(Vision), Gemini Pro Vision, Claude 3, and LLaVA-1.5, by emphasizing nuanced understanding and interpretation of visual data. The benchmark comprises 346 images paired with 1129 questions, all meticulously crafted by human experts. We introduce a novel structure for these visual questions designed to establish control groups. This structure enables us to conduct a quantitative analysis of the models' response tendencies, logical consistency, and various failure modes. In our evaluation on HallusionBench, we benchmarked 15 different models, highlighting a 31.42% question-pair accuracy achieved by the state-of-the-art GPT-4V. Notably, all other evaluated models achieve accuracy below 16%. Moreover, our analysis not only highlights the observed failure modes, including language hallucination and visual illusion, but also deepens an understanding of these pitfalls. Our comprehensive case studies within HallusionBench shed light on the challenges of hallucination and illusion in LVLMs. Based on these insights, we suggest potential pathways for their future improvement. The benchmark and codebase can be accessed at https://github.com/tianyi-lab/HallusionBench.

📄 PDF Abstract BibTeX arXiv:2310.14566

Code (9)

tianyi-lab/hallusionbench 공식 구현
FuxiaoLiu/LRV-Instruction pytorch
FuxiaoLiu/VisualNews-Repository pytorch
codelion/adaptive-classifier pytorch
dongping-chen/mllm-as-a-judge pytorch
fuxiaoliu/mmc pytorch
wuxiyang1996/AutoHallusion pytorch
zli12321/qa_metrics
zli12321/videohallu pytorch

Tasks

DiagnosticHallucinationHallucination EvaluationVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models

2026-05-03 · Zeshang Li, Shuoyang Zhang arxiv

Vision-Language Models (VLMs) hallucinate objects that are not present, and a growing line of work tries to curb this by feeding the model its own generated caption as auxiliary evidence -- assuming that a caption, once …

Improving Explainability of Disentangled Representations using Multipath-Attribution Mappings

2023-06-15 · Lukas Klein, João B. S. Carvalho, Mennatallah El-Assady, Paolo Penna 외

Explainable AI aims to render model behavior understandable by humans, which can be seen as an intermediate step in extracting causal relations from correlative patterns. Due to the high risk of possible fatal decisions …

Relation Extraction

Beyond Dominant Patches: Spatial Credit Redistribution For Grounded Vision-Language Models

2026-02-25 · Niamul Hassan Samin, Md Arifur Rahman, Abdullah Ibne Hanif Arean, Juena Ahmed Noshin 외 arxiv

Vision-Language Models (VLMs) often hallucinate objects that are not present in the input image. We identify a contributing cause of this behavior, which we term spatial credit collapse: in early transformer layers, hidd…

An Advanced NLP Framework for Automated Medical Diagnosis with DeBERTa and Dynamic Contextual Positional Gating

2025-02-11 · Mohammad Ali Labbaf Khaniki, Sahabeh Saadati, Mohammad Manthouri

This paper presents a novel Natural Language Processing (NLP) framework for enhancing medical diagnosis through the integration of advanced techniques in data augmentation, feature extraction, and classification. The pro…

ClassificationData AugmentationDiagnosticMedical Diagnosis

What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models

2019-07-31 · TACL 2020 1 · Allyson Ettinger

Pre-training by language modeling has become a popular and successful approach to NLP tasks, but we have yet to understand exactly what linguistic capacities these pre-training processes confer upon models. In this paper…

Language ModelingLanguage ModellingNegation