paper-with-me

홈 › Papers

Measuring and Aligning Abstraction in Vision-Language Models with Medical Taxonomies

2026-01-21 · Ben Schaper, Maxime Di Folco, Bernhard Kainz, Julia A. Schnabel, Cosmin I. Bercea arxiv

Vision-Language Models show strong zero-shot performance for chest X-ray classification, but standard flat metrics fail to distinguish between clinically minor and severe errors. This work investigates how to quantify and mitigate abstraction errors by leveraging medical taxonomies. We benchmark several state-of-the-art VLMs using hierarchical metrics and introduce Catastrophic Abstraction Errors to capture cross-branch mistakes. Our results reveal substantial misalignment of VLMs with clinical taxonomies despite high flat performance. To address this, we propose risk-constrained thresholding and taxonomy-aware fine-tuning with radial embeddings, which reduce severe abstraction errors to below 2 per cent while maintaining competitive performance. These findings highlight the importance of hierarchical evaluation and representation-level alignment for safer and more clinically meaningful deployment of VLMs.

📄 PDF Abstract BibTeX arXiv:2601.14827

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visual Instruction Tuning Aligns Modalities through Abstraction

2026-06-02 · Luis Palacios, Lorenzo Basile, Diego Doimo, Alberto Cazzaniga arxiv

Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text. Yet, it remains unclear how visual features are embedded into the layer-wise hierarchy of…

DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

2024-05-31 · Linli Yao, Lei LI, Shuhuai Ren, Lean Wang 외

The visual projector, which bridges the vision and language modalities and facilitates cross-modal alignment, serves as a crucial component in MLLMs. However, measuring the effectiveness of projectors in vision-language …

cross-modal alignmentVisual LocalizationVisual Question Answering (VQA)

When Does Perceptual Alignment Benefit Vision Representations?

2024-10-14 · Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Y. Tamir 외

Humans judge perceptual similarity according to diverse visual attributes, including scene layout, subject location, and camera pose. Existing vision models understand a wide range of semantic abstractions but improperly…

Depth EstimationImage GenerationInductive BiasRetrieval+1

Beyond Literal Summarization: Redefining Hallucination for Medical SOAP Note Evaluation

2026-04-16 · Bhavik Vachhani, Kush Shrisvastava, Pranshu Nema, Sai Chiranthan arxiv

Evaluating large language models (LLMs) for clinical documentation tasks such as SOAP note generation remains challenging. Unlike standard summarization, these tasks require clinical abstraction, normalization of colloqu…

MAIN-VLA: Modeling Abstraction of Intention and eNvironment for Vision-Language-Action Models

2026-02-02 · Zheyuan Zhou, Liang Du, Zixun Sun, Xiaoyu Zhou 외 arxiv

Despite significant progress in Visual-Language-Action (VLA), in highly complex and dynamic environments that involve real-time unpredictable interactions (such as 3D open worlds and large-scale PvP games), existing appr…