paper-with-me

홈 › Papers

MedFM-Robust: Benchmarking Robustness of Medical Foundation Models

2026-05-18 · Xiangxiang Cui, Tianjin Huang, Yifang Wang, Lijie Hu, Lu Yin arxiv

Medical foundation models have achieved remarkable clinical performance, yet their robustness under real-world perturbations remains underexplored. We present a robustness benchmark comprising 40 perturbation types (12 base, 28 medical-specific) across eight imaging modalities, evaluating five VLMs (LLaVA-Med, MedGemma, MedGemma-1.5, Gemini-2.5-flash and GPT-4o-mini) on VQA, visual grounding, and captioning, alongside two segmentation models (MedSAM, SAM-Med2D) with five fine-tuning strategies. Our findings reveal: (1) Fine-tuning strategy dominates robustness, with LoRA exhibiting nearly double the degradation of full fine-tuning, while SAM-Med2D's Adapter offers favorable efficiency-robustness trade-off. (2) Medical-specific perturbations disproportionately damage segmentation, with 9 of 15 top corruptions being domain-specific. (3) LoRA-tuned visual grounding drops over 40 points, whereas zero-shot captioning remains stable (<7% drop). Zero-shot VQA shows model-dependent robustness--medical models drop under 20% while Gemini-2.5-flash drops 54%. General-purpose VLMs achieve higher VQA accuracy but fail on grounding; among medical VLMs, MedGemma demonstrates the best overall stability. These results provide deployment guidelines and underscore the necessity of domain-specific robustness evaluation for medical AI. Our code is available at: https://abnerai.github.io/MedFM-Robust.

📄 PDF Abstract BibTeX arXiv:2605.19027

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

FairMedFM: Fairness Benchmarking for Medical Imaging Foundation Models

2024-07-01 · Ruinan Jin, Zikang Xu, Yuan Zhong, Qiongsong Yao 외

The advent of foundation models (FMs) in healthcare offers unprecedented opportunities to enhance medical diagnostics through automated classification and segmentation tasks. However, these models also raise significant …

BenchmarkingFairnessparameter-efficient fine-tuningZero-Shot Learning

TransMed: Large Language Models Enhance Vision Transformer for Biomedical Image Classification

2023-12-12 · Kaipeng Zheng, Weiran Huang, Lichao Sun

Few-shot learning has been studied to adapt models to tasks with very few samples. It holds profound significance, particularly in clinical tasks, due to the high annotation cost of medical images. Several works have exp…

Few-Shot Learningimage-classificationImage Classification

MedFMC: A Real-world Dataset and Benchmark For Foundation Model Adaptation in Medical Image Classification

2023-06-16 · Dequan Wang, Xiaosong Wang, Lilong Wang, Mengzhang Li 외

Foundation models, often pre-trained with large-scale data, have achieved paramount success in jump-starting various vision and language applications. Recent advances further enable adapting foundation models in downstre…

Diabetic Retinopathy Gradingimage-classificationImage ClassificationIn-Context Learning+3

An Investigation of Visual Foundation Models Robustness

2025-08-22 · Sandeep Gupta, Roberto Passerone arxiv

Visual Foundation Models (VFMs) are becoming ubiquitous in computer vision, powering systems for diverse tasks such as object detection, image classification, segmentation, pose estimation, and motion tracking. VFMs are …

Image ClassificationObject DetectionPose Estimation

MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare

2026-05-12 · Yihao Wang, Haoran Xu, Renjie Gu, Yixuan Ye 외 arxiv

The large-scale deployment of personalized healthcare agents demands memory mechanisms that are exceptionally precise, safe, and capable of long-term clinical tracking. However, existing benchmarks primarily focus on dai…