paper-with-me

Papers

THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models

2024-05-08 · CVPR 2024 1 · Prannay Kaul, Zhizhong Li, Hao Yang, Yonatan Dukler, Ashwin Swaminathan, C. J. Taylor, Stefano Soatto

Mitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which we term "Type I hallucinations". Instead, they focus on hallucinations responding to very specific question formats -- typically a multiple-choice response regarding a particular object or attribute -- which we term "Type II hallucinations". Additionally, such benchmarks often require external API calls to models which are subject to change. In practice, we observe that a reduction in Type II hallucinations does not lead to a reduction in Type I hallucinations but rather that the two forms of hallucinations are often anti-correlated. To address this, we propose THRONE, a novel object-based automatic framework for quantitatively evaluating Type I hallucinations in LVLM free-form outputs. We use public language models (LMs) to identify hallucinations in LVLM responses and compute informative metrics. By evaluating a large selection of recent LVLMs using public datasets, we show that an improvement in existing metrics do not lead to a reduction in Type I hallucinations, and that established benchmarks for measuring Type I hallucinations are incomplete. Finally, we provide a simple and effective data augmentation method to reduce Type I and Type II hallucinations as a strong baseline.

📄 PDF Abstract BibTeX arXiv:2405.05256

Code (1)

haoyu-bu/CAFe pytorch

Tasks

AttributeData AugmentationFormHallucinationMultiple-choice

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

From Uncertainty to Trust: Enhancing Reliability in Vision-Language Models with Uncertainty-Guided Dropout Decoding

2024-12-09 · Yixiong Fang, Ziran Yang, Zhaorun Chen, Zhuokai Zhao 외

Large vision-language models (LVLMs) demonstrate remarkable capabilities in multimodal tasks but are prone to misinterpreting visual inputs, often resulting in hallucinations and unreliable outputs. To address these chal…

Who Wins the Game of Thrones? How Sentiments Improve the Prediction of Candidate Choice

2020-02-29 · Chaehan So

This paper analyzes how candidate choice prediction improves by different psychological predictors. To investigate this question, it collected an original survey dataset featuring the popular TV series "Game of Thrones".…

BenchmarkingHoldout Set

GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity

2025-08-27 · Seongheon Park, Sharon Li arxiv

Object hallucination in large vision-language models presents a significant challenge to their safe deployment in real-world applications. Recent works have proposed object-level hallucination scores to estimate the like…

Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models

2024-10-30 · Junjie Wu, Tsz Ting Chung, Kai Chen, Dit-yan Yeung

Despite the outstanding performance in vision-language reasoning, Large Vision-Language Models (LVLMs) might generate hallucinated contents that do not exist in the given image. Most existing LVLM hallucination benchmark…

HallucinationHallucination EvaluationObjectObject Hallucination+2

Comparison between Automatic and Human Subtitling: A Case Study with Game of Thrones

2019-09-01 · RANLP 2019 9 · Sabrina Baldo de Br{\'e}bisson

In this submission, I would like to share my experiences with the software DeepL and the comparison analysis I have made with human subtitling offered by the DVD version of the corpus I have chosen as the topic of my stu…

Translation