paper-with-me

홈 › Papers

RAU: Reference-based Anatomical Understanding with Vision Language Models

2025-09-26 · Yiwei Li, Yikang Liu, Jiaqi Guo, Lin Zhao, Zheyuan Zhang, Xiao Chen, Boris Mailhe, Ankush Mukherjee, Terrence Chen, Shanhui Sun arxiv

Anatomical understanding through deep learning is critical for automatic report generation, intra-operative navigation, and organ localization in medical imaging; however, its progress is constrained by the scarcity of expert-labeled data. A promising remedy is to leverage an annotated reference image to guide the interpretation of an unlabeled target. Although recent vision-language models (VLMs) exhibit non-trivial visual reasoning, their reference-based understanding and fine-grained localization remain limited. We introduce RAU, a framework for reference-based anatomical understanding with VLMs. We first show that a VLM learns to identify anatomical regions through relative spatial reasoning between reference and target images, trained on a moderately sized dataset. We validate this capability through visual question answering (VQA) and bounding box prediction. Next, we demonstrate that the VLM-derived spatial cues can be seamlessly integrated with the fine-grained segmentation capability of SAM2, enabling localization and pixel-level segmentation of small anatomical regions, such as vessel segments. Across two in-distribution and two out-of-distribution datasets, RAU consistently outperforms a SAM2 fine-tuning baseline using the same memory setup, yielding more accurate segmentations and more reliable localization. More importantly, its strong generalization ability makes it scalable to out-of-distribution datasets, a property crucial for medical image applications. To the best of our knowledge, RAU is the first to explore the capability of VLMs for reference-based identification, localization, and segmentation of anatomical structures in medical images. Its promising performance highlights the potential of VLM-driven approaches for anatomical understanding in automated clinical workflows.

📄 PDF Abstract BibTeX arXiv:2509.22404

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringSpatial ReasoningVisual Reasoning

Similar Papers 제목 키워드 기반

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

2026-06-07 · Sergios Gatidis, Curtis Langlotz, Christian Bluethgen arxiv

Vision-language models (VLMs) pretrained on large-scale image-text pairs demonstrate strong image-level understanding, but are primarily optimized for global alignment and do not explicitly encode fine-grained anatomical…

AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

2026-01-06 · Anees Ur Rehman Hashmi, Numan Saeed, Christoph Lippert arxiv

Multimodal medical large language models have shown substantial progress in chest X-ray interpretation but continue to face challenges in spatial reasoning and anatomical understanding. Although existing grounding techni…

Visual Question AnsweringSpatial ReasoningPhrase Grounding

U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding

2025-05-23 · Anjie Le, Henan Liu, Yue Wang, Zhenyu Liu 외

Ultrasound is a widely-used imaging modality critical to global healthcare, yet its interpretation remains challenging due to its varying image quality on operators, noises, and anatomical structures. Although large visi…

BenchmarkingSpatial ReasoningText Generation

Your other Left! Vision-Language Models Fail to Identify Relative Positions in Medical Images

2025-08-01 · Daniel Wolf, Heiko Hillenhagen, Billurvan Taskin, Alex Bäuerle 외 arxiv

Clinical decision-making relies heavily on understanding relative positions of anatomical structures and anomalies. Therefore, for Vision-Language Models (VLMs) to be applicable in clinical practice, the ability to accur…

Towards Anatomically Plausible Human Image Generation via Synthetic Localized Preferences

2026-05-25 · Bao Li, Yuliang Xiu, Zhen Liu arxiv

Large-scale text-to-image foundation models have achieved remarkable visual realism, yet generating human images with correct anatomical structures remains challenging. Existing approaches enforce anatomical constraints …

Image Generation