paper-with-me

홈 › Papers

Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation

2026-04-02 · Ji Young Byun, Young-Jin Park, Jean-Philippe Corbeil, Asma Ben Abacha arxiv

As vision-language models (VLMs) are increasingly deployed in clinical decision support, more than accuracy is required: knowing when to trust their predictions is equally critical. Yet, a comprehensive and systematic investigation into the overconfidence of these models remains notably scarce in the medical domain. We address this gap through a comprehensive empirical study of confidence calibration in VLMs, spanning three model families (Qwen3-VL, InternVL3, LLaVA-NeXT), three model scales (2B--38B), and multiple confidence estimation prompting strategies, across three medical visual question answering (VQA) benchmarks. Our study yields three key findings: First, overconfidence persists across model families and is not resolved by scaling or prompting, such as chain-of-thought and verbalized confidence variants. Second, simple post-hoc calibration approaches, such as Platt scaling, reduce calibration error and consistently outperform the prompt-based strategy. Third, due to their (strict) monotonicity, these post-hoc calibration methods are inherently limited in improving the discriminative quality of predictions, leaving AUROC at the same level. Motivated by these findings, we investigate hallucination-aware calibration (HAC), which incorporates vision-grounded hallucination detection signals as complementary inputs to refine confidence estimates. We find that leveraging these hallucination signals improves both calibration and AUROC, with the largest gains on open-ended questions. Overall, our findings suggest post-hoc calibration as standard practice for medical VLM deployment over raw confidence estimates, and highlight the practical usefulness of hallucination signals to enable more reliable use of VLMs in medical VQA.

📄 PDF Abstract BibTeX arXiv:2604.02543

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models

2025-02-16 · Prateek Chhikara

Large Language Models (LLMs) demonstrate impressive performance across diverse tasks, yet confidence calibration remains a challenge. Miscalibration - where models are overconfident or underconfident - poses risks, parti…

Multiple-choice

Confidence Calibration in Large Language Model-Based Entity Matching

2025-09-23 · Iris Kamsteeg, Juan Cardenas-Cartagena, Floris van Beers, Gineke ten Holt 외 arxiv

This research aims to explore the intersection of Large Language Models and confidence calibration in Entity Matching. To this end, we perform an empirical study to compare baseline RoBERTa confidences for an Entity Matc…

Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of Overconfidence

2021-12-01 · NeurIPS 2021 12 · Deng-Bao Wang, Lei Feng, Min-Ling Zhang

Capturing accurate uncertainty quantification of the prediction from deep neural networks is important in many real-world decision-making applications. A reliable predictor is expected to be accurate when it is confident…

Decision MakingUncertainty Quantification

The Dunning-Kruger Effect in Large Language Models: An Empirical Study of Confidence Calibration

2026-02-12 · Sudipta Ghosh, Mrityunjoy Panday arxiv

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet their ability to accurately assess their own confidence remains poorly understood. We present an empirical study investigat…

Beyond Overconfidence: Foundation Models Redefine Calibration in Deep Neural Networks

2025-06-11 · Achim Hekler, Lukas Kuhn, Florian Buettner

Reliable uncertainty calibration is essential for safely deploying deep neural networks in high-stakes applications. Deep neural networks are known to exhibit systematic overconfidence, especially under distribution shif…