VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence
Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimodal reasoning, yet current multimodal large language model (MLLM) evaluations do not explicitly enforce this alignment. Existing benchmarks assess either textual answer correctness or pixel-level localization in isolation, leaving the coupling of reasoning and grounding an open challenge. We introduce VISTAQA, a comprehensive benchmark for joint evaluation of free-form answer correctness and pixel-level evidence grounding in visual question answering. VISTAQA comprises 1,157 expert-curated samples spanning six task types and six visual domains, ranging from direct perception to compositional and relational reasoning. VISTAQA requires models to not only answer correctly, but to also provide precise segmentation masks that support their answers. It also includes hallucination-aware examples where no valid visual evidence exists. To support this enhanced evaluation, we introduce GROVE, a unified evaluation metric that enforces joint correctness by combining textual accuracy and grounding quality via a per-sample geometric mean, ensuring neither dimension can compensate for deficiencies in the other. Comprehensive experiments across grounding-aware models and hybrid pipelines with general-purpose MLLMs reveal that even the strongest systems achieve limited performance under GROVE, highlighting a substantial gap between answer accuracy and visual evidence alignment.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual Question AnsweringRelational ReasoningMultimodal ReasoningSimilar Papers 제목 키워드 기반
Latent Variable Models for Visual Question Answering
Current work on Visual Question Answering (VQA) explore deterministic approaches conditioned on various types of image and question features. We posit that, in addition to image and question pairs, other modalities are u…
BenchmarkingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA systems relied heavily on language biases, motivating subsequent work …
Visual Question AnsweringRepresentation LearningMachine TranslationImage CaptioningWhat's Different between Visual Question Answering for Machine "Understanding" Versus for Accessibility?
In visual question answering (VQA), a machine must answer a question given an associated image. Recently, accessibility researchers have explored whether VQA can be deployed in a real-world setting where users with visua…
BenchmarkingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Chart Question Answering from Real-World Analytical Narratives
We present a new dataset for chart question answering (CQA) constructed from visualization notebooks. The dataset features real-world, multi-view charts paired with natural language questions grounded in analytical narra…
Chart Question AnsweringBenchmarking Geospatial Question Answering Engines using the Dataset GeoQuestions1089
We present the dataset GeoQuestions1089 for benchmarking geospatial question answering engines. GeoQuestions1089 is the largest such dataset available presently and it contains 1089 questions, their corresponding GeoSPA…
BenchmarkingKnowledge Base Question AnsweringQuestion Answering