paper-with-me

홈 › Papers

Confidence in the Reasoning of Large Language Models

2024-12-19 · Yudi Pawitan, Chris Holmes

There is a growing literature on reasoning by large language models (LLMs), but the discussion on the uncertainty in their responses is still lacking. Our aim is to assess the extent of confidence that LLMs have in their answers and how it correlates with accuracy. Confidence is measured (i) qualitatively in terms of persistence in keeping their answer when prompted to reconsider, and (ii) quantitatively in terms of self-reported confidence score. We investigate the performance of three LLMs -- GPT4o, GPT4-turbo and Mistral -- on two benchmark sets of questions on causal judgement and formal fallacies and a set of probability and statistical puzzles and paradoxes. Although the LLMs show significantly better performance than random guessing, there is a wide variability in their tendency to change their initial answers. There is a positive correlation between qualitative confidence and accuracy, but the overall accuracy for the second answer is often worse than for the first answer. There is a strong tendency to overstate the self-reported confidence score. Confidence is only partially explained by the underlying token-level probability. The material effects of prompting on qualitative confidence and the strong tendency for overconfidence indicate that current LLMs do not have any internally coherent sense of confidence.

📄 PDF Abstract BibTeX arXiv:2412.15296

Code (1)

yudpaw-git/statspuzzle 공식 구현

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning

2025-09-16 · Caiqi Zhang, Chang Shu, Ehsan Shareghi, Nigel Collier arxiv

Confidence estimation is essential for the reliable deployment of large language models (LLMs). Existing methods are primarily designed for factual QA tasks and often fail to generalize to reasoning tasks. To address thi…

Confidence-Calibrated Small-Large Language Model Collaboration for Cost-Efficient Reasoning

2026-03-04 · Chuang Zhang, Zizhen Zhu, Yihao Wei, Bing Tian 외 arxiv

Large language models (LLMs) demonstrate superior reasoning capabilities compared to small language models (SLMs), but incur substantially higher costs. We propose COllaborative REAsoner (COREA), a system that cascades a…

Reinforcement Learning

Confidence over Time: Confidence Calibration with Temporal Logic for Large Language Model Reasoning

2026-01-19 · Zhenjiang Mao, Anirudhh Venkat, Artem Bisliouk, Akshat Kothiyal 외 arxiv

Large Language Models (LLMs) increasingly rely on long-form, multi-step reasoning to solve complex tasks such as mathematical problem solving and scientific question answering. Despite strong performance, existing confid…

Question Answering

Confidence Geometry Reveals Trace-Level Correctness in Large Language Model Reasoning

2026-05-16 · Shuo Liu, Ding Liu, Shi-Ju Ran arxiv

Large language models (LLMs) generate not only reasoning text, but also token-level confidence trajectories that record how uncertainty evolves during inference. Whether these trajectories are relevant to reasoning corre…

VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning

2026-04-10 · Wenyi Xiao, Xinchi Xu, Leilei Gan arxiv

Large Vision Language Models (LVLMs) achieve strong multimodal reasoning but frequently exhibit hallucinations and incorrect responses with high certainty, which hinders their usage in high-stakes domains. Existing verba…

Reinforcement LearningMultimodal ReasoningVisual ReasoningVisual Grounding