paper-with-me

홈 › Papers

Confidence Calibration in Large Language Models

2026-04-03 · Noam Michael, Daniel BenShushan, Jacob Bien, Don A. Moore arxiv

We investigate the calibration of large language models' (LLMs') confidence across diverse tasks. The results of our preregistered study show that the current crop of LLMs are, like people, too sure they are right: confidence exceeds accuracy, on average. Importantly, however, this tendency is moderated by a powerful hard-easy effect, wherein overconfidence is greatest on difficult tests; by contrast, easy tests actually show substantial underconfidence. We develop LifeEval, a test for evaluating model calibration across levels of difficulty.

📄 PDF Abstract BibTeX arXiv:2605.23909

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Confidence Calibration in Large Language Model-Based Entity Matching

2025-09-23 · Iris Kamsteeg, Juan Cardenas-Cartagena, Floris van Beers, Gineke ten Holt 외 arxiv

This research aims to explore the intersection of Large Language Models and confidence calibration in Entity Matching. To this end, we perform an empirical study to compare baseline RoBERTa confidences for an Entity Matc…

Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models

2025-02-16 · Prateek Chhikara

Large Language Models (LLMs) demonstrate impressive performance across diverse tasks, yet confidence calibration remains a challenge. Miscalibration - where models are overconfident or underconfident - poses risks, parti…

Multiple-choice

CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?

2025-10-28 · Qing Zong, Jiayu Liu, Tianshi Zheng, Chunyang Li 외 arxiv

Accurate confidence calibration in Large Language Models (LLMs) is critical for safe use in high-stakes domains, where clear verbalized confidence enhances user trust. Traditional methods that mimic reference confidence …

Calibrating the Confidence of Large Language Models by Eliciting Fidelity

2024-04-03 · Mozhi Zhang, Mianqiu Huang, Rundong Shi, Linsen Guo 외

Large language models optimized with techniques like RLHF have achieved good alignment in being helpful and harmless. However, post-alignment, these language models often exhibit overconfidence, where the expressed confi…

Language ModelingLanguage Modelling

VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning

2026-04-10 · Wenyi Xiao, Xinchi Xu, Leilei Gan arxiv

Large Vision Language Models (LVLMs) achieve strong multimodal reasoning but frequently exhibit hallucinations and incorrect responses with high certainty, which hinders their usage in high-stakes domains. Existing verba…

Reinforcement LearningMultimodal ReasoningVisual ReasoningVisual Grounding