paper-with-me

홈 › Papers

Enhancing Medical Visual Grounding via Knowledge-guided Spatial Prompts

2026-04-02 · Yifan Gao, Tao Zhou, Yi Zhou, Ke Zou, Yizhe Zhang, Huazhu Fu arxiv

Medical Visual Grounding (MVG) aims to identify diagnostically relevant phrases from free-text radiology reports and localize their corresponding regions in medical images, providing interpretable visual evidence to support clinical decision-making. Although recent Vision-Language Models (VLMs) exhibit promising multimodal reasoning ability, their grounding remains insufficient spatial precision, largely due to a lack of explicit localization priors when relying solely on latent embeddings. In this work, we analyze this limitation from an attention perspective and propose KnowMVG, a Knowledge-prior and global-local attention enhancement framework for MVG in VLMs that explicitly strengthens spatial awareness during decoding. Specifically, we present a knowledge-enhanced prompting strategy that encodes phrase related medical knowledge into compact embeddings, together with a global-local attention that jointly leverages coarse global information and refined local cues to guide precise region localization. localization. This design bridges high-level semantic understanding and fine-grained visual perception without introducing extra textual reasoning overhead. Extensive experiments on four MVG benchmarks demonstrate that our KnowMVG consistently outperforms existing approaches, achieving gains of 3.0% in AP50 and 2.6% in mIoU over prior state-of-the-art methods. Qualitative and ablation studies further validate the effectiveness of each component.

📄 PDF Abstract BibTeX arXiv:2604.01915

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningVisual Grounding

Similar Papers 제목 키워드 기반

Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions

2025-03-05 · Jun Li, Che Liu, Wenjia Bai, Rossella Arcucci 외

Visual Language Models (VLMs) have demonstrated impressive capabilities in visual grounding tasks. However, their effectiveness in the medical domain, particularly for abnormality detection and localization within medica…

Anomaly DetectionVisual Grounding

MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive Learning

2024-02-03 · CVPR 2024 1 · Zhe Li, Laurence T. Yang, Bocheng Ren, Xin Nie 외

The scarcity of annotated data has sparked significant interest in unsupervised pre-training methods that leverage medical reports as auxiliary signals for medical visual representation learning. However, existing resear…

Contrastive Learningimage-classificationImage Classificationobject-detection+4

XMedFusion: A Knowledge-Guided Multimodal Perception and Reasoning Framework for Autonomous Medical Systems

2026-06-08 · Hamza Riaz, Arham Haroon, Maha Baig, Muhammad Dawood Rizwan 외 arxiv

Autonomous medical and robotic systems increasingly rely on intelligent perception and reasoning capabilities to interpret visual data and support clinical decision making. Radiology report generation represents a critic…

Visual GroundingDecision Making

MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training in Radiology

2023-01-05 · Chaoyi Wu, Xiaoman Zhang, Ya zhang, Yanfeng Wang 외

In this paper, we consider enhancing medical visual-language pre-training (VLP) with domain-specific knowledge, by exploiting the paired image-text reports from the radiological daily practice. In particular, we make the…

Medical DiagnosisSelf-Supervised LearningTriplet

MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray Diagnosis

2023-01-01 · ICCV 2023 1 · Chaoyi Wu, Xiaoman Zhang, Ya zhang, Yanfeng Wang 외

In this paper, we consider enhancing medical visual-language pre-training (VLP) with domain-specific knowledge, by exploiting the paired image-text reports from the radiological daily practice. In particular, we make…

Medical DiagnosisTriplet