paper-with-me

홈 › Papers

Pathological Truth Bias in Vision-Language Models

2025-09-14 · Yash Thube arxiv

Vision Language Models (VLMs) are improving quickly, but standard benchmarks can hide systematic failures that reduce real world trust. We introduce MATS (Multimodal Audit for Truthful Spatialization), a compact behavioral audit that measures whether models reject visually contradicted statements, and two metrics Spatial Consistency Score (SCS) and Incorrect Agreement Rate (IAR). Instruction tuned generative VLMs (LLaVA 1.5, QwenVLchat) exhibit very low SCS and high IAR, while contrastive encoders (CLIP, SigLIP) are far more robust. Activation patching causally localizes failure loci (mid to late cross attention for generative models, pooled projection components for contrastive models) and suggests concrete repair paths.

📄 PDF Abstract BibTeX arXiv:2509.22674

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Re-thinking and Re-labeling LIDC-IDRI for Robust Pulmonary Cancer Prediction

2022-07-28 · Hanxiao Zhang, Xiao Gu, Minghui Zhang, Weihao Yu 외

The LIDC-IDRI database is the most popular benchmark for lung cancer prediction. However, with subjective assessment from radiologists, nodules in LIDC may have entirely different malignancy annotations from the patholog…

Metric LearningPredictionRetrieval

Shape-aware synthesis of pathological lung CT scans using CycleGAN for enhanced semi-supervised lung segmentation

2024-05-14 · Rezkellah Noureddine Khiati, Pierre-Yves Brillet, Aurélien Justet, Radu Ispas 외

This paper addresses the problem of pathological lung segmentation, a significant challenge in medical image analysis, particularly pronounced in cases of peripheral opacities (severe fibrosis and consolidation) because …

Data AugmentationImage SegmentationImage-to-Image TranslationMedical Image Analysis+3

CXR-LLAVA: a multimodal large language model for interpreting chest X-ray images

2023-10-22 · Seowoo Lee, Jiwon Youn, Hyungjin Kim, Mansu Kim 외

Purpose: This study aimed to develop an open-source multimodal large language model (CXR-LLAVA) for interpreting chest X-ray images (CXRs), leveraging recent advances in large language models (LLMs) to potentially replic…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model+3

LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation

2026-02-04 · Ruixiao Yang, Yuanhe Tian, Xu Yang, Huiqi Li 외 arxiv

Radiology Report Generation (RRG) aims to produce accurate and coherent diagnostics from medical images. Although large vision language models (LVLM) improve report fluency and accuracy, they exhibit hallucinations, gene…

Libra-MIL: Multimodal Prototypes Stereoscopic Infused with Task-specific Language Priors for Few-shot Whole Slide Image Classification

2025-11-11 · Zhenfeng Zhuang, Fangyu Zhou, Liansheng Wang arxiv

While Large Language Models (LLMs) are emerging as a promising direction in computational pathology, the substantial computational cost of giga-pixel Whole Slide Images (WSIs) necessitates the use of Multi-Instance Learn…

Image Classification