paper-with-me

Papers

6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models

2025-12-03 · Leon Mayer, Piotr Kalinowski, Caroline Ebersbach, Marcel Knopp, Tim Rädsch, Evangelia Christodoulou, Annika Reinke, Fiona R. Kolbinger, Lena Maier-Hein arxiv

Vision-language models (VLMs) are increasingly integrated into clinical workflows. However, existing benchmarks primarily assess performance on common anatomical presentations and fail to capture the challenges posed by rare variants. To address this gap, we introduce AdversarialAnatomyBench, the first benchmark comprising naturally occurring rare anatomical variants across diverse imaging modalities and anatomical regions. We call such variants that violate learned priors about "typical" human anatomy natural adversarial anatomy. Benchmarking 25 state-of-the-art VLMs with AdversarialAnatomyBench yielded three key insights. First, when queried with basic medical perception tasks, mean accuracy dropped from 71% on typical to 28% on atypical anatomy. Even the best-performing models, GPT-5, Gemini 2.5 Pro, and Llama 4 Maverick, showed performance drops of 41-51%. Second, model errors closely mirrored expected anatomical biases. Third, neither model scaling nor interventions, including bias-aware prompting and test-time reasoning, resolved these issues. These findings highlight a critical limitation in current VLMs: their poor generalization to rare anatomical presentations. AdversarialAnatomyBench provides a foundation for systematically measuring and mitigating anatomical bias in multimodal medical artificial intelligence (AI) systems.

📄 PDF Abstract BibTeX arXiv:2512.04238

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Terabyte-scale supervised 3D training and benchmarking dataset of the mouse kidney

2021-08-04 · Willy Kuo, Diego Rossinelli, Georg Schulz, Roland H. Wenger 외

The performance of machine learning algorithms, when used for segmenting 3D biomedical images, does not reach the level expected based on results achieved with 2D photos. This may be explained by the comparative lack of …

BenchmarkingBIG-bench Machine LearningData AugmentationTransfer Learning

Unsupervised Medical Image Segmentation with Adversarial Networks: From Edge Diagrams to Segmentation Maps

2019-11-12 · Umaseh Sivanesan, Luis H. Braga, Ranil R. Sonnadara, Kiret Dhindsa

We develop and approach to unsupervised semantic medical image segmentation that extends previous work with generative adversarial networks. We use existing edge detection methods to construct simple edge diagrams, train…

Edge DetectionImage SegmentationMedical Image SegmentationSegmentation+1

Kidney and Kidney Tumour Segmentation in CT Images

2022-12-26 · Qi Ming How, Hoi Leong Lee

Automatic segmentation of kidney and kidney tumour in Computed Tomography (CT) images is essential, as it uses less time as compared to the current gold standard of manual segmentation. However, many hospitals are still …

Computed Tomography (CT)Segmentation

Hierarchical Perceptual Noise Injection for Social Media Fingerprint Privacy Protection

2022-08-23 · Simin Li, Huangxinxin Xu, Jiakai Wang, Aishan Liu 외

Billions of people are sharing their daily life images on social media every day. However, their biometric information (e.g., fingerprint) could be easily stolen from these images. The threat of fingerprint leakage from …

Adversarial Attack

FedAgain: A Trust-Based and Robust Federated Learning Strategy for an Automated Kidney Stone Identification in Ureteroscopy

2026-03-19 · Ivan Reyes-Amezcua, Francisco Lopez-Tiro, Clément Larose, Christian Daul 외 arxiv

The reliability of artificial intelligence (AI) in medical imaging critically depends on its robustness to heterogeneous and corrupted images acquired with diverse devices across different hospitals which is highly chall…

Federated Learning