paper-with-me

홈 › Papers

Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations

2025-03-18 · Ziwei Ji, Lei Yu, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Cheng Zhang, Pascale Fung, Nicola Cancedda

LLMs often adopt an assertive language style also when making false claims. Such `overconfident hallucinations'' mislead users and erode trust. Achieving the ability to express in language the actual degree of uncertainty around a claim is therefore of great importance. We find that verbal uncertainty'' is governed by a single linear feature in the representation space of LLMs, and show that this has only moderate correlation with the actual `semantic uncertainty'' of the model. We apply this insight and show that (1) the mismatch between semantic and verbal uncertainty is a better predictor of hallucinations than semantic uncertainty alone and (2) we can intervene on verbal uncertainty at inference time and reduce hallucinations on short-form answers, achieving an average relative reduction of 32%.

📄 PDF Abstract BibTeX arXiv:2503.14477

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation

2025-12-23 · Bhaktipriya Radharapu, Eshika Saxena, Kenneth Li, Chenxi Whitehouse 외 arxiv

As LLM-based judges become integral to industry applications, obtaining well-calibrated uncertainty estimates efficiently has become critical for production deployment. However, existing techniques, such as verbalized co…

Calibrating Verbalized Probabilities for Large Language Models

2024-10-09 · Cheng Wang, Gyuri Szarvas, Georges Balazs, Pavel Danchenko 외

Calibrating verbalized probabilities presents a novel approach for reliably assessing and leveraging outputs from black-box Large Language Models (LLMs). Recent methods have demonstrated improved calibration by applying …

I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation

2026-04-05 · Haotian Zong, Binze Li, Yufei Long, Sinyin Chang 외 arxiv

Large language models (LLMs) frequently produce confident but incorrect answers, partly because common binary scoring conventions reward answering over honestly expressing uncertainty. We study whether prompt-only interv…

Calibrating LLMs with Semantic-level Reward

2026-05-15 · Fengfei Yu, Ruijia Niu, Dongxia Wu, Yian Ma 외 arxiv

As large language models (LLMs) are deployed in consequential settings such as medical question answering and legal reasoning, the ability to estimate when their outputs are likely to be correct is essential for safe and…

Reinforcement LearningQuestion AnsweringLegal Reasoning

SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment

2026-06-09 · Kaiwen Shi, Zheyuan Zhang, Yanfang Ye arxiv

Large language models increasingly express uncertainty through natural-language statements, yet these expressions often fail to reflect the model's sampled behavior. We study verbal uncertainty alignment as a distributio…