paper-with-me

홈 › Papers

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting

2026-06-30 · Qian Ma, S M Rayeed, Charles V. Stewart, Qiong Wu, Yao Ma arxiv

Knowledge-Based Visual Question Answering (KB-VQA) aims to evaluate whether Visual Language Models (VLMs) can retrieve, ground, and reason over external structured knowledge beyond visual evidence. In practice, answer accuracy is widely adopted as the primary evaluation metric, implicitly treating correctness as a proxy for knowledge-grounded reasoning. However, for existing KB-VQA benchmarks, this proxy relies on critical assumptions that are often overlooked and rendered unreliable by benchmark issues: annotated answer must be derivable from the associated knowledge base, question must be well-posed with sufficient constraints, and visual setting must meaningfully require grounded disambiguation. In this work, we show that these assumptions are systematically violated in existing KB-VQA benchmarks. Our audit reveals substantial instances with missing or contradicted answers and underspecified questions that render accuracy a misleading metric. Furthermore, we find that existing datasets rely on visually trivial, single-entity scenes that bypass the need for sophisticated visual-to-knowledge mapping. We demonstrate that even with controlled architectures, these flaws lead to distorted model rankings and overestimations of reasoning capabilities. To address this, we introduce (1) a principled audit-and-repair protocol that restores answer derivability and question clarity, and (2) a controlled multi-entity augmentation protocol that introduces visual ambiguity to challenge initial retrieval and grounded reasoning. Re-evaluation under corrected and augmented settings yields markedly different performance trends. Our findings call for rethinking evaluation protocols and designing more interaction-aware KB-VQA benchmarks that prioritize verifiable reasoning over simple matching.

📄 PDF Abstract BibTeX arXiv:2607.00159

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets

2021-08-01 · ACL 2021 5 · Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim 외

Auditing NLP systems for computational harms like surfacing stereotypes is an elusive goal. Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompan…

coreference-resolutionCoreference ResolutionFairnessLanguage Modeling+1

Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia

2025-06-10 · Katelyn Xiaoying Mei, Anna Seo Gyeong Choi, Hilke Schellmann, Mona Sloane 외

Automatic Speech Recognition (ASR) has transformed daily tasks from video transcription to workplace hiring. ASR systems' growing use warrants robust and standardized auditing approaches to ensure automated transcription…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Navigatespeech-recognition+1

Bound by the Bounty: Collaboratively Shaping Evaluation Processes for Queer AI Harms

2023-07-15 · Organizers Of QueerInAI, Nathan Dennler, Anaelia Ovalle, Ashwin Singh 외

Bias evaluation benchmarks and dataset and model documentation have emerged as central processes for assessing the biases and harms of artificial intelligence (AI) systems. However, these auditing processes have been cri…

Adaptive Plan-Execute Framework for Smart Contract Security Auditing

2025-05-21 · Zhiyuan Wei, Jing Sun, Zijian Zhang, Zhe Hou 외

Large Language Models (LLMs) have shown great promise in code analysis and auditing; however, they still struggle with hallucinations and limited context-aware reasoning. We introduce SmartAuditFlow, a novel Plan-Execute…

RAGRetrieval-augmented GenerationVulnerability Detection

PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning

2026-06-16 · Bo Su, Ankit Shah, Thai Le arxiv

Machine unlearning for large language models (LLMs) aims to remove specified knowledge while preserving the rest of the model's capabilities. However, the boundary between knowledge to forget and knowledge to retain is o…