paper-with-me

Papers

VisionTrap: Unanswerable Questions On Visual Data

2025-07-23 · Asir Saadat, Syem Aziz, Shahriar Mahmud, Abdullah Ibne Masud Mahi, Sabbir Ahmed arxiv

Visual Question Answering (VQA) has been a widely studied topic, with extensive research focusing on how VLMs respond to answerable questions based on real-world images. However, there has been limited exploration of how these models handle unanswerable questions, particularly in cases where they should abstain from providing a response. This research investigates VQA performance on unrealistically generated images or asking unanswerable questions, assessing whether models recognize the limitations of their knowledge or attempt to generate incorrect answers. We introduced a dataset, VisionTrap, comprising three categories of unanswerable questions across diverse image types: (1) hybrid entities that fuse objects and animals, (2) objects depicted in unconventional or impossible scenarios, and (3) fictional or non-existent figures. The questions posed are logically structured yet inherently unanswerable, testing whether models can correctly recognize their limitations. Our findings highlight the importance of incorporating such questions into VQA benchmarks to evaluate whether models tend to answer, even when they should abstain.

📄 PDF Abstract BibTeX arXiv:2507.17262

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Unanswerable Questions about Images and Texts

2021-01-25 · Ernest Davis

Questions about a text or an image that cannot be answered raise distinctive issues for an AI. This note discusses the problem of unanswerable questions in VQA (visual question answering), in QA (visual question answerin…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions

2024-10-05 · Xingwei He, Qianru Zhang, A-Long Jin, Yuan Yuan 외

Large Vision-Language Models (LVLMs) have achieved remarkable progress on visual perception and linguistic interpretation. Despite their impressive capabilities across various tasks, LVLMs still suffer from the issue of …

BenchmarkingHallucinationMathematical ReasoningMME+3

Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents

2025-11-14 · Davide Napolitano, Luca Cagliero, Fabrizio Battiloro arxiv

The evolution of Visual Large Language Models (VLLMs) has revolutionized the automatic understanding of Visually Rich Documents (VRDs), which contain both textual and visual elements. Although VLLMs excel in Visual Quest…

Visual Question Answering

CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering

2025-01-02 · Ben Vardi, Oron Nir, Ariel Shamir

Recent Vision-Language Models (VLMs) have demonstrated remarkable capabilities in visual understanding and reasoning, and in particular on multiple-choice Visual Question Answering (VQA). Still, these models can make dis…

Multiple-choiceQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Long-Form Answers to Visual Questions from Blind and Low Vision People

2024-08-12 · Mina Huh, Fangyuan Xu, Yi-Hao Peng, Chongyan Chen 외

Vision language models can now generate long-form answers to questions about images - long-form visual question answers (LFVQA). We contribute VizWiz-LF, a dataset of long-form answers to visual questions posed by blind …

FormVisual Question Answering (VQA)