paper-with-me

Papers

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study

2025-07-15 · Che Liu, Jiazhen Pan, Weixiang Shen, Wenjia Bai, Daniel Rueckert, Rossella Arcucci arxiv

Vision-Language Models (VLMs) trained on web-scale corpora excel at natural image tasks and are increasingly repurposed for healthcare; however, their competence in medical tasks remains underexplored. We present a comprehensive evaluation of open-source general-purpose and medically specialised VLMs, ranging from 3B to 72B parameters, across eight benchmarks: MedXpert, OmniMedVQA, PMC-VQA, PathVQA, MMMU, SLAKE, and VQA-RAD. To observe model performance across different aspects, we first separate it into understanding and reasoning components. Three salient findings emerge. First, large general-purpose models already match or surpass medical-specific counterparts on several benchmarks, demonstrating strong zero-shot transfer from natural to medical images. Second, reasoning performance is consistently lower than understanding, highlighting a critical barrier to safe decision support. Third, performance varies widely across benchmarks, reflecting differences in task design, annotation quality, and knowledge demands. No model yet reaches the reliability threshold for clinical deployment, underscoring the need for stronger multimodal alignment and more rigorous, fine-grained evaluation protocols.

📄 PDF Abstract BibTeX arXiv:2507.11200

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-modal Pre-training for Medical Vision-language Understanding and Generation: An Empirical Study with A New Benchmark

2023-06-10 · Li Xu, Bo Liu, Ameer Hamza Khan, Lu Fan 외

With the availability of large-scale, comprehensive, and general-purpose vision-language (VL) datasets such as MSCOCO, vision-language pre-training (VLP) has become an active area of research and proven to be effective f…

Image-text RetrievalMedical Report GenerationQuestion AnsweringRetrieval+2

Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey

2025-07-09 · Getamesay Haile Dagnaw, Yanming Zhu, Muhammad Hassan Maqsood, Wencheng Yang 외

Explainable artificial intelligence (XAI) has become increasingly important in biomedical image analysis to promote transparency, trust, and clinical adoption of DL models. While several surveys have reviewed XAI techniq…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)Survey

A Survey of Medical Vision-and-Language Applications and Their Techniques

2024-11-19 · Qi Chen, Ruoshan Zhao, Sinuo Wang, Vu Minh Hieu Phan 외

Medical vision-and-language models (MVLMs) have attracted substantial interest due to their capability to offer a natural language interface for interpreting complex medical data. Their applications are versatile and hav…

Decision MakingDiagnosticImage-text RetrievalMedical Report Generation+5

Comprehensive language-image pre-training for 3D medical image understanding

2025-10-16 · Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor 외 arxiv

Vision-language pre-training, i.e., aligning images with paired text, is a powerful paradigm to create encoders that can be directly used for tasks such as classification, retrieval, and segmentation. In the 3D medical i…

Semantic Segmentation

MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output

2025-10-11 · Yanyuan Chen, Dexuan Xu, Yu Huang, Songkun Zhan 외 arxiv

Currently, medical vision language models are widely used in medical vision question answering tasks. However, existing models are confronted with two issues: for input, the model only relies on text instructions and lac…

Instruction FollowingQuestion Answering