paper-with-me

홈 › Papers

Large Language Models Are Overconfident in Their Own Responses

2026-06-02 · Mario Sanz-Guerrero, Manuel Mager, Katharina von der Wense arxiv

Prior work has shown that instruction-tuned large language models (LLMs) are less well calibrated than their base pre-trained counterparts. However, little is known about the frequently used chat template's effect on the calibration of conversational LLMs. In this work, we investigate the mechanisms driving this miscalibration by decoupling the effects of the post-training algorithm and the chat format. We find that, while instruction tuning fundamentally harms calibration, the chat template aggravates the issue through an "ownership bias" -- models are significantly more confident in their own answers than in identical answers provided by a user. Extensive experiments across six recent open-weight LLMs, three benchmarks, and three confidence elicitation methods show that models assign up to 26% higher confidence to their own responses. Leveraging this insight, we propose a simple inference-time strategy: framing the model's answer as user input during confidence elicitation. This approach significantly reduces overconfidence and improves calibration by up to 26% without the need for retraining, narrowing the gap between base and instruction-tuned models.

📄 PDF Abstract BibTeX arXiv:2606.03437

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Humans overrely on overconfident language models, across languages

2025-07-08 · Neil Rathi, Dan Jurafsky, Kaitlyn Zhou arxiv

As large language models (LLMs) are deployed globally, it is crucial that their responses are calibrated across languages to accurately convey uncertainty and limitations. Prior work shows that LLMs are linguistically ov…

Identifying Influential N-grams in Confidence Calibration via Regression Analysis

2026-04-07 · Shintaro Ozaki, Wataru Hashimoto, Hidetaka Kamigaito, Katsuhiko Hayashi 외 arxiv

While large language models (LLMs) improve performance by explicit reasoning, their responses are often overconfident, even though they include linguistic expressions demonstrating uncertainty. In this work, we identify …

Reconfidencing LLMs from the Grouping Loss Perspective

2024-02-07 · Lihu Chen, Alexandre Perez-Lebel, Fabian M. Suchanek, Gaël Varoquaux

Large Language Models (LLMs), including ChatGPT and LLaMA, are susceptible to generating hallucinated answers in a confident tone. While efforts to elicit and calibrate confidence scores have proven useful, recent findin…

Uncertainty Quantification

INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection

2024-02-06 · Chao Chen, Kai Liu, Ze Chen, Yi Gu 외

Knowledge hallucination have raised widespread concerns for the security and reliability of deployed LLMs. Previous efforts in detecting hallucinations have been employed at logit-level uncertainty estimation or language…

DiversityHallucinationQuestion Answering

Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents

2026-06-01 · Aitor Arronte Alvarez, Naiyi Xie Fincham arxiv

Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly used in these systems to provide scalable, personalized feedback. Howe…

Bias Detection