paper-with-me

Papers

DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams

2026-04-28 · Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Gaurav Najpande, Manan Suri, Dinesh Manocha, Puneet Mathur, Vivek Gupta arxiv

Diagram question answering (DQA) requires models to interpret structured visual representations such as charts, maps, infographics, circuit schematics, and scientific diagrams. Recent vision-language models (VLMs) often achieve high answer accuracy on these tasks, yet correct answers do not guarantee that models ground their reasoning in the diagram regions that support the prediction. Models may instead rely on textual correlations or dataset artifacts without identifying the visual evidence required to verify the answer. This limitation prevents reliable evaluation of diagram reasoning and reduces interpretability. We introduce DRAGON, a benchmark for evaluating evidence-grounded visual reasoning in diagrams. Given a diagram, a question, and the correct answer, a model must predict bounding boxes that correspond to the visual elements required to justify the answer. These evidence regions may include answer-bearing components, textual labels, legends, axes, connectors, and other supporting structures involved in the reasoning process. The DRAGON dataset contains 11,664 annotated question instances collected from six diagram QA datasets: ChartQA, Circuit-VQA, InfographicsVQA, MapIQ, MapWise, and AI2D. We release a 2,445-instance benchmark test set with human-verified reasoning evidence annotations and a standardized evaluation framework. We evaluate eight recent VLMs and analyze their ability to localize reasoning evidence across diverse diagram domains. DRAGON enables systematic evaluation of diagram reasoning and supports future research on models that ground their predictions in visual evidence.

📄 PDF Abstract BibTeX arXiv:2604.25231

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Reasoning

Similar Papers 제목 키워드 기반

Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology

2025-07-10 · Haochen Wang, Xiangtai Li, Zilong Huang, Anran Wang 외 arxiv

Models like OpenAI-o3 pioneer visual grounded reasoning by dynamically referencing visual regions, just like human "thinking with images". However, no benchmark exists to evaluate these capabilities holistically. To brid…

Reinforcement LearningObject Localization

CXReasonAgent: Evidence-Grounded Diagnostic Reasoning Agent for Chest X-rays

2026-02-26 · Hyungyung Lee, Hangyul Yoon, Edward Choi arxiv

Chest X-ray plays a central role in thoracic diagnosis, and its interpretation inherently requires multi-step, evidence-grounded reasoning. However, large vision-language models (LVLMs) often generate plausible responses…

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

2026-07-13 · Wencheng Ye, Yi Bin, Yujuan Ding, Hongye Fang 외 arxiv

Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning. To understand thi…

Multimodal Reasoning

Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs

2026-05-23 · Wen Ma, Fucheng Niu, Zhiting Fan, Zikai Xiao 외 arxiv

Vision-language models have demonstrated impressive capabilities in general medical visual question answering, yet due to limited interpretability, it remains unclear whether their predictions reflect evidence-grounded c…

Visual Question AnsweringAdversarial RobustnessVisual Grounding

Bridging Modality Disconnect in Self-Reflection via Closed-Loop Visually Grounded Verification

2026-02-21 · Haoyu Zhang, Yuwei Wu, Pengxiang Li, Xintong Zhang 외 arxiv

In the era of Vision-Language Models (VLMs), enhancing multimodal reasoning capabilities remains a critical challenge, particularly in handling ambiguous or complex visual inputs, where initial inferences often lead to h…

Multimodal Reasoning