paper-with-me

홈 › Papers

Predictive Entropy Links Calibration and Paraphrase Sensitivity in Medical Vision-Language Models

2026-04-10 · Binesh Sadanandan, Vahid Behzadan arxiv

Medical Vision Language Models VLMs suffer from two failure modes that threaten safe deployment mis calibrated confidence and sensitivity to question rephrasing. We show they share a common cause, proximity to the decision boundary, by benchmarking five uncertainty quantification methods on MedGemma 4BIT across in distribution MIMIC CXR and outof distribution PadChest chest X ray datasets, with cross architecture validation on LLaVA RAD7B. For well calibrated single model methods, predictive entropy from one forward pass predicts which samples will flip under rephrasing AUROC 0.711 on MedGemma, 0.878 on LLaVARAD p 10 4, enabling a single entropy threshold to flag both unreliable and rephrase sensitive predictions. A five member LoRA ensemble fails under the MIMIC PadChest shift 42.9 ECE, 34.1 accuracy, though LLaVA RAD s ensemble does not collapse 69.1. MC Dropout achieves the best calibration ECE 4.3 and selective prediction coverage 21.5 at 5 risk, yet total entropy from a single forward pass outperforms the ensemble for both error detection AUROC 0.743 vs 0.657 and paraphrase screening. Simple methods win.

📄 PDF Abstract BibTeX arXiv:2604.08941

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology

2026-05-18 · Franciskus Xaverius Erick, Johanna Paula Müller, Bernhard Kainz arxiv

Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational histopathology. However, fine-tuning billions of parameters on scarce, expert-annota…

D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models

2025-10-10 · Jisu Han, Wonjun Hwang arxiv

Test-time adaptation paradigm provides flexibility towards domain shifts by performing immediate adaptation on unlabeled target data from the source model. Vision-Language Models (VLMs) leverage their generalization capa…

Test-time Adaptation

Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models

2026-08-06 · Jingyan Jiang, Yaru Sun, Xiao Chen, Jiazhen Huang 외 arxiv

Test-time adaptation (TTA) can improve the recognition accuracy of vision-language models under distribution shift, but often degrades calibration, making predictive confidence unreliable for downstream decision-making. …

Test-time Adaptation

Guarding the Meaning: Self-Supervised Training for Semantic Robustness in Guard Models

2025-11-06 · Cristina Pinneri, Christos Louizos arxiv

Guard models are a critical component of LLM safety, but their sensitivity to superficial linguistic variations remains a key vulnerability. We show that even meaning-preserving paraphrases can cause large fluctuations i…

CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation

2026-08-13 · Enhan Li, Junhao He, Hongyang Du arxiv

On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this l…