paper-with-me

홈 › Papers

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

2026-07-02 · Sofiane Ouaari, Kevin Vorwalder, Nico Pfeifer arxiv

Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answering. Medical Image Quality Assessment (MIQA) supports diagnostic accuracy and patient safety by determining whether images meet the standards required for clinical decision-making. Automating MIQA with VLMs may reduce workload, but their behavior under real-world conditions, where images may be degraded or textual context may affect judgments, should be further explored before deployment. We benchmark VLMs on medical image quality using the MediMeta-C dataset zero-shot across seven corruption types and five severity levels. We evaluate sensitivity to degradation patterns, the effect of corruptions on embedding geometry, and whether textual attributes (demographics, expertise, infrastructure, institution) alter scores. Across 16 VLMs and seven modalities, pixelation produced the largest score reductions (mean -20.58%, up to -34.4% for OCT), whereas brightness had limited effect (-0.81%). Embedding displacement was associated with score changes. Same-family models showed correlations of 0.67-0.83; some produced increases up to +31% for corrupted mammography. Textual attributes affected scores: institutional prestige raised them +17.15%, and equipment age lowered them -14.7%. The largest changes were +95.62% (InternVL-8B) and -37.7% (MedGemma). Current VLMs show limitations for medical image quality assessment. Pixelation, a privacy-preserving transformation, reduces performance, indicating a trade-off between patient privacy and reliability. Sensitivity to contextual metadata indicates limited objectivity and marks metadata as a privacy and bias source. Privacy protection and objective quality assessment are related requirements for use.

📄 PDF Abstract BibTeX arXiv:2607.01973

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringImage Quality Assessment

Similar Papers 제목 키워드 기반

Feature Extraction for Generative Medical Imaging Evaluation: New Evidence Against an Evolving Trend

2023-11-22 · McKell Woodland, Austin Castelo, Mais Al Taie, Jessica Albuquerque Marques Silva 외

Fr\'echet Inception Distance (FID) is a widely used metric for assessing synthetic image quality. It relies on an ImageNet-based feature extractor, making its applicability to medical imaging unclear. A recent trend is t…

Data AugmentationMedical Image Generation

A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis

2023-10-31 · Yingshu Li, Yunyi Liu, Zhanyu Wang, Xinyu Liang 외

This work conducts an evaluation of GPT-4V's multimodal capability for medical image analysis, with a focus on three representative tasks of radiology report generation, medical visual question answering, and medical vis…

DescriptiveMedical Image AnalysisMedical Visual Question AnsweringQuestion Answering+3

Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs

2024-09-24 · Shadi Iskander, Nachshon Cohen, Zohar Karnin, Ori Shapira 외

Training large language models (LLMs) for external tool usage is a rapidly expanding field, with recent research focusing on generating synthetic data to address the shortage of available data. However, the absence of sy…

Assessing Intra-class Diversity and Quality of Synthetically Generated Images in a Biomedical and Non-biomedical Setting

2023-07-23 · Muhammad Muneeb Saad, Mubashir Husain Rehmani, Ruairi O'Reilly

In biomedical image analysis, data imbalance is common across several imaging modalities. Data augmentation is one of the key solutions in addressing this limitation. Generative Adversarial Networks (GANs) are increasing…

Data AugmentationDiversity

Enhancing Reliability of Medical Image Diagnosis through Top-rank Learning with Rejection Module

2025-08-11 · Xiaotong Ji, Ryoma Bise, Seiichi Uchida arxiv

In medical image processing, accurate diagnosis is of paramount importance. Leveraging machine learning techniques, particularly top-rank learning, shows significant promise by focusing on the most crucial instances. How…