paper-with-me

Papers

H-POPE: Hierarchical Polling-based Probing Evaluation of Hallucinations in Large Vision-Language Models

2024-11-06 · Nhi Pham, Michael Schott

By leveraging both texts and images, large vision language models (LVLMs) have shown significant progress in various multi-modal tasks. Nevertheless, these models often suffer from hallucinations, e.g., they exhibit inconsistencies between the visual input and the textual output. To address this, we propose H-POPE, a coarse-to-fine-grained benchmark that systematically assesses hallucination in object existence and attributes. Our evaluation shows that models are prone to hallucinations on object existence, and even more so on fine-grained attributes. We further investigate whether these models rely on visual input to formulate the output texts.

📄 PDF Abstract BibTeX arXiv:2411.04077

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationObject

Similar Papers 제목 키워드 기반

What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?

2025-08-03 · Ming-Kun Xie, Jia-Hao Xiao, Gang Niu, Lei Feng 외 arxiv

Large Vision-Language Models (LVLMs), empowered by the success of Large Language Models (LLMs), have achieved impressive performance across domains. Despite the great advances in LVLMs, they still suffer from the unavail…

Instruction Makes a Difference

2024-02-01 · Tosin Adewumi, Nudrat Habib, Lama Alkhaled, Elisa Barney

We introduce Instruction Document Visual Question Answering (iDocVQA) dataset and Large Language Document (LLaDoc) model, for training Language-Vision (LV) models for document analysis and predictions on document images,…

HallucinationInstruction FollowingObject HallucinationQuestion Answering+1

Evaluating Object Hallucination in Large Vision-Language Models

2023-05-17 · YiFan Li, Yifan Du, Kun Zhou, Jinpeng Wang 외

Inspired by the superior language abilities of large language models (LLM), large vision-language models (LVLM) have been recently explored by integrating powerful LLMs for improving the performance on complex multimodal…

HallucinationObjectObject Hallucination

Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions

2025-09-19 · Hansol Park, Hoseong Ahn, Junwon Moon, Yejin Lee 외 arxiv

Hallucinations in multimodal models have been extensively studied using benchmarks that probe reliability in image-text query settings. However, the effect of spoken queries on multimodal hallucinations remains largely u…

DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark

2024-11-05 · Haodong Li, Haicheng Qu, Xiaofeng Zhang

With the rapid development of large vision language models (LVLMs), these models have shown excellent results in various multimodal tasks. Since LVLMs are prone to hallucinations and there are currently few datasets and …

Data AugmentationHallucinationHallucination Evaluation