paper-with-me

홈 › Papers

Beyond Accuracy: Evaluating Visual Grounding In Multimodal Medical Reasoning

2026-03-03 · Anas Zafar, Leema Krishna Murali, Ashish Vashist arxiv

Recent work shows that text-only reinforcement learning with verifiable rewards (RLVR) can match or outperform image-text RLVR on multimodal medical VQA benchmarks, suggesting current evaluation protocols may fail to measure causal visual dependence. We introduce a counterfactual evaluation framework using real, blank, and shuffled images across four medical VQA benchmarks: PathVQA, PMC-VQA, SLAKE, and VQA-RAD. Beyond accuracy, we measure Visual Reliance Score (VRS), Image Sensitivity (IS), and introduce Hallucinated Visual Reasoning Rate (HVRR) to detect cases where models generate visual claims despite producing image-invariant answers. Our findings reveal that RLVR improves accuracy while degrading visual grounding: text-only RLVR achieves negative VRS on PathVQA (-0.09), performing better with mismatched images, while image-text RLVR reduces image sensitivity to 39.8% overall despite improving accuracy. On VQA-RAD, both variants achieve 63% accuracy through different mechanisms: text-only RLVR retains 81% performance with blank images, while image-text RLVR shows only 29% image sensitivity. Models generate visual claims in 68-74% of responses, yet 38-43% are ungrounded (HVRR). These findings demonstrate that accuracy-only rewards enable shortcut exploitation, and progress requires grounding-aware evaluation protocols and training objectives that explicitly enforce visual dependence.

📄 PDF Abstract BibTeX arXiv:2603.03437

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningVisual ReasoningVisual Grounding

Similar Papers 제목 키워드 기반

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

2025-06-05 · Tianxu Wang, Zhuofan Zhang, Ziyu Zhu, Yue Fan 외

3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplored. In this paper, we introduce Anywhere3…

3D visual groundingObjectReferring ExpressionSpatial Reasoning+1

Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision

2026-03-27 · Ling Li, Bowen Liu, Zinuo Zhan, Peng Jie 외 arxiv

Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores non-verbal deictic cues prevalent in re…

Referring ExpressionVisual Grounding

M$^3$Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

2026-06-05 · Zhengjun Huang, Wenxuan Liu, Zhoujin Tian, Wei Chen 외 arxiv

Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visuals and straightforward content, evaluating neither reasoning over aut…

Multimodal Fact-Level Attribution for Verifiable Reasoning

2026-02-12 · David Wan, Han Wang, Ziyang Wang, Elias Stengel-Eskin 외 arxiv

Multimodal large language models (MLLMs) are increasingly used for real-world tasks involving multi-step reasoning and long-form generation, where reliability requires grounding model outputs in heterogeneous input sourc…

Multimodal Reasoning

\textsc{GUI-Spotlight}: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding

2025-10-05 · Bin Lei, Nuo Xu, Ali Payani, Mingyi Hong 외 arxiv

Multimodal large language models (MLLMs) have markedly expanded the competence of graphical user-interface (GUI) systems, propelling them beyond controlled simulations into complex, real-world environments across diverse…

Visual Grounding