paper-with-me

홈 › Papers

Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation

2026-03-02 · Aditya Parikh, Aasa Feragen, Sneha Das, Stella Frank arxiv

Reliable deployment of Vision-Language Models (VLMs) in radiology requires validation metrics that go beyond surface-level text similarity to ensure clinical fidelity and demographic fairness. This paper investigates a critical blind spot in current model evaluation: the use of decoding strategies that lead to high aggregate token-overlap scores despite succumbing to template collapse, in which models generate only repetitive, safe generic text and omit clinical terminology. Unaddressed, this blind spot can lead to metric gaming, where models that perform well on benchmarks prove clinically uninformative. Instead, we advocate for lexical diversity measures to check model generations for clinical specificity. We introduce Clinical Association Displacement (CAD), a vocabulary-level framework that quantifies shifts in demographic-based word associations in generated reports. Weighted Association Erasure (WAE) aggregates these shifts to measure the clinical signal loss across demographic groups. We show that deterministic decoding produces high levels of semantic erasure, while stochastic sampling generates diverse outputs but risks introducing new bias, motivating a fundamental rethink of how "optimal" reporting is defined.

📄 PDF Abstract BibTeX arXiv:2603.01625

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pathological Truth Bias in Vision-Language Models

2025-09-14 · Yash Thube arxiv

Vision Language Models (VLMs) are improving quickly, but standard benchmarks can hide systematic failures that reduce real world trust. We introduce MATS (Multimodal Audit for Truthful Spatialization), a compact behavior…

What sentiment analysis can't see: Measuring whether customers were helped, and what went wrong, across 70,000 support conversations

2026-06-18 · Jason Potteiger arxiv

Most companies read their customer support data at scale using sentiment analysis, which measures how customers sound rather than whether they were satisfied with the result. We tested a richer alternative on 70,450 supp…

Sentiment Analysis

Measuring Visual Understanding in Telecom domain: Performance Metrics for Image-to-UML conversion using VLMs

2025-09-15 · HG Ranjani, Rutuja Prabhudesai arxiv

Telecom domain 3GPP documents are replete with images containing sequence diagrams. Advances in Vision-Language Large Models (VLMs) have eased conversion of such images to machine-readable PlantUML (puml) formats. Howeve…

Measuring and Aligning Abstraction in Vision-Language Models with Medical Taxonomies

2026-01-21 · Ben Schaper, Maxime Di Folco, Bernhard Kainz, Julia A. Schnabel 외 arxiv

Vision-Language Models show strong zero-shot performance for chest X-ray classification, but standard flat metrics fail to distinguish between clinically minor and severe errors. This work investigates how to quantify an…

Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM

2024-04-29 · Navid Rajabi, Jana Kosecka

Vision and Language Models (VLMs) continue to demonstrate remarkable zero-shot (ZS) performance across various tasks. However, many probing studies have revealed that even the best-performing VLMs struggle to capture asp…

Phrase GroundingScene Understanding