paper-with-me

홈 › Papers

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs

2026-07-04 · Anas Zafar, Leema Krishna Murali, Siddhant Bharadwaj, Ashish Vashist, Jia Wu arxiv

Large vision language models (VLMs) report strong accuracy on medical question-answering, yet it remains unclear whether they reason from visual evidence or exploit textual shortcuts. We introduce a counterfactual evaluation framework that decouples visual and textual contributions by substituting input images with controlled surrogates blank, pixel-shuffled, image-absent, and CLIP-retrieved hard negatives and derive a suite of grounding metrics including the Visual Reliance Score (VRS) and Visual Hallucination Rate (VHR). We further introduce CORAL (COntrastive Retrieval-Augmented Learning), a 7B-parameter LoRA fine-tune of Qwen2.5-VL-7B trained with a Contrastive Grounding Objective (CGO) that penalises answer invariance under hard-negative image swaps. On a paired controlled evaluation across four closed-form medical VQA benchmarks (PathVQA, PMC-VQA, SLAKE, VQA-RAD; n=400 total), CORAL improves macro accuracy by +6.7 pp (P(Delta>0)=0.988) and reduces VHR by 8.0 pp (P<0.001) over the matched Qwen2.5-VL-7B base; neither MedVLThinker RL variant achieves a significant gain on either metric. Cross-domain diagnostics further reveal that image substitution costs only <=6.5 pp on medical benchmarks versus 48-61 pp on general-domain tasks, situating the grounding gap that CGO targets. We discuss evaluation limitations openly including train/eval benchmark overlap and underpowered secondary metrics and release our framework, training code, and model weights to support reproducible grounding audits of medical VLMs.

📄 PDF Abstract BibTeX arXiv:2607.03647

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting Clinical Hallucinations in LVLMs via Counterfactual Visual Grounding Uncertainty

2026-06-26 · Xiao Song, Haonan Qin, Zhaoxu Zhang, Jiong Zhang 외 arxiv

Large vision-language models (LVLMs) are increasingly used for clinical image understanding, yet they remain vulnerable to \emph{hallucinations}--producing textual findings or attributes not supported by the image. We pr…

Visual Grounding

Modularized Textual Grounding for Counterfactual Resilience

2019-04-07 · CVPR 2019 6 · Zhiyuan Fang, Shu Kong, Charless Fowlkes, Yezhou Yang

Computer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries. To achieve high grounding precision, current textual grounding meth…

AttributecounterfactualNatural Language Visual GroundingPhrase Grounding+2

Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language Grounding

2020-12-01 · NeurIPS 2020 12 · Zhu Zhang, Zhou Zhao, Zhijie Lin, Jieming Zhu 외

Weakly-supervised vision-language grounding aims to localize a target moment in a video or a specific region in an image according to the given sentence query, where only video-level or image-level sentence annotations a…

Contrastive LearningcounterfactualRelationSentence

HalluSegBench: Counterfactual Visual Reasoning for Segmentation Hallucination Evaluation

2025-06-26 · Xinzhuo Li, Adheesh Juvekar, Xingyou Liu, Muntasir Wahed 외

Recent progress in vision-language segmentation has significantly advanced grounded visual understanding. However, these models often exhibit hallucinations by producing segmentation masks for objects not grounded in the…

counterfactualCounterfactual ReasoningHallucinationHallucination Evaluation+4

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

2025-08-19 · Catherine Glossop, William Chen, Arjun Bhorkar, Dhruv Shah 외 arxiv

Generalist robots should be able to understand and follow user instructions. Despite providing a powerful architecture for mapping open-vocabulary language instructions to robot actions, current vision-language-action (V…

Vision-Language NavigationInstruction Following