paper-with-me

Papers

Confidence Estimation for LLMs in Multi-turn Interactions

2026-01-05 · Caiqi Zhang, Ruihan Yang, Xiaochen Zhu, Chengzu Li, Tiancheng Hu, Yijiang River Dong, Deqing Yang, Nigel Collier arxiv

While confidence estimation is a promising direction for mitigating hallucinations in Large Language Models (LLMs), current research overwhelmingly focuses on single-turn settings. The dynamics of model confidence in multi-turn conversations, where context accumulates and ambiguity is progressively resolved, remain largely unexplored. This work presents the first systematic study of confidence estimation in multi-turn interactions, establishing a formal evaluation framework grounded in two key desiderata: per-turn calibration and monotonicity of confidence as more information becomes available. To facilitate this, we introduce novel metrics, including a length-normalized Expected Calibration Error (InfoECE), and a new "Hinter-Guesser" paradigm for generating controlled evaluation datasets. Our experiments reveal that widely-used confidence techniques struggle with calibration and monotonicity in multi-turn dialogues. In contrast, a novel logit-based probe we introduce, P(Sufficient), proves comparatively more effective, robustly tracking evidence accumulation and distinguishing it from conversational filler. Our work provides a foundational methodology for developing more reliable and trustworthy conversational agents.

📄 PDF Abstract BibTeX arXiv:2601.02179

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Confidence Should Be Calibrated More Than One Turn Deep

2026-04-07 · Zhaohan Zhang, Chengzhengxu Li, Xiaoming Liu, Chao Shen 외 arxiv

Large Language Models (LLMs) are increasingly applied in high-stakes domains such as finance, healthcare, and education, where reliable multi-turn interactions with users are essential. However, existing work on confiden…

BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents

2025-10-27 · Litu Ou, Kuan Li, Huifeng Yin, Liwen Zhang 외 arxiv

Confidence in LLMs is a useful indicator of model uncertainty and answer reliability. Existing work mainly focused on single-turn scenarios, while research on confidence in complex multi-turn interactions is limited. In …

When Two LLMs Debate, Both Think They'll Win

2025-05-25 · Pradyumna Shyama Prasad, Minh Nhat Nguyen

Can LLMs accurately adjust their confidence when facing opposition? Building on previous studies measuring calibration on static fact-based question-answering tasks, we evaluate Large Language Models (LLMs) in a dynamic,…

Question Answering

CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models

2025-12-16 · Yiran Zhang, Jincheng Hu, Mark Dras, Usman Naseem arxiv

Large language models (LLMs) excel at single-turn reasoning but often lose accuracy and coherence over extended, multi-turn interactions. Recent evaluations such as TurnBench highlight recurring failure modes-reasoning b…

High-Confidence Off-Policy (or Counterfactual) Variance Estimation

2021-01-25 · Yash Chandak, Shiv Shankar, Philip S. Thomas

Many sequential decision-making systems leverage data collected using prior policies to propose a new policy. For critical applications, it is important that high-confidence guarantees on the new policy's behavior are pr…

counterfactualDecision MakingSequential Decision MakingVocal Bursts Intensity Prediction