paper-with-me

홈 › Papers

Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models

2023-10-09 · Holy Lovenia, Wenliang Dai, Samuel Cahyawijaya, Ziwei Ji, Pascale Fung

Object hallucination poses a significant challenge in vision-language (VL) models, often leading to the generation of nonsensical or unfaithful responses with non-existent objects. However, the absence of a general measurement for evaluating object hallucination in VL models has hindered our understanding and ability to mitigate this issue. In this work, we present NOPE (Negative Object Presence Evaluation), a novel benchmark designed to assess object hallucination in VL models through visual question answering (VQA). We propose a cost-effective and scalable approach utilizing large language models to generate 29.5k synthetic negative pronoun (NegP) data of high quality for NOPE. We extensively investigate the performance of 10 state-of-the-art VL models in discerning the non-existence of objects in visual questions, where the ground truth answers are denoted as NegP (e.g., "none"). Additionally, we evaluate their standard performance on visual questions on 9 other VQA datasets. Through our experiments, we demonstrate that no VL model is immune to the vulnerability of object hallucination, as all models achieve accuracy below 10\% on NegP. Furthermore, we uncover that lexically diverse visual questions, question types with large scopes, and scene-relevant objects capitalize the risk of object hallucination in VL models.

📄 PDF Abstract BibTeX arXiv:2310.05338

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationObjectObject HallucinationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

NoPE: The Counting Power of Transformers with No Positional Encodings

2025-05-16 · Chris Köcher, Alexander Kozachinskiy, Anthony Widjaja Lin, Marco Sälzer 외

Positional Encodings (PEs) seem to be indispensable for ensuring expressiveness of transformers; without them attention transformers reduce to a bag-of-word model. NoPE-transformers (i.e. with No PEs) with unique hard at…

Hard Attention

Development and validation of an artificial intelligence model to accurately predict spinopelvic parameters

2024-02-09 · Edward S. Harake, Joseph R. Linzey, Cheng Jiang, Rushikesh S. Joshi 외

Objective. Achieving appropriate spinopelvic alignment has been shown to be associated with improved clinical symptoms. However, measurement of spinopelvic radiographic parameters is time-intensive and interobserver reli…

DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios

2025-03-25 · Xiangting Meng, Jiaqi Yang, Mingshu Chen, Chenxin Yan 외

In the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent. However, the scarcity of corresponding real-world datasets significantly hinders the development and eval…

3D Object DetectionObjectobject-detectionObject Tracking+2

Rethinking Efficient Graph Coarsening via a Non-Selfishness Principle

2026-05-13 · Xu Bai, Bin Lu, Kun Zhang, Shengbo Chen 외 arxiv

Graph coarsening is a graph dimensionality reduction technique that aims to construct a smaller and more tractable graph while preserving the essential structural and semantic properties of the original graph. However, m…

Dimensionality Reduction

NOPE: Novel Object Pose Estimation from a Single Image

2023-03-23 · CVPR 2024 1 · Van Nguyen Nguyen, Thibault Groueix, Yinlin Hu, Mathieu Salzmann 외

The practicality of 3D object pose estimation remains limited for many applications due to the need for prior knowledge of a 3D model and a training period for new objects. To address this limitation, we propose an appro…

ObjectPose Estimation