paper-with-me

홈 › Papers

Early-Token Confidence Predicts Reasoning Quality in Multi-Agent LLM Debate

2026-06-09 · Ali Keramati, Justin Cheok, Jacob Horne, Mark Warschauer arxiv

Evaluating reasoning quality in multi-agent LLM systems is challenging, especially for open-ended tasks without reference answers. We investigate whether intrinsic confidence signals, token-level log-probabilities from decoding, can predict reasoning quality as assessed by LLM-as-judge evaluation. Using a debate-based essay scoring framework, we compare confidence proxies against rubric-based judge scores across two ASAP essay sets. We find that early-token confidence, particularly within the first few generated tokens, is consistently the strongest predictor of reasoning quality, outperforming full-sequence statistics. Analysis of log-probability trajectories shows that the opening phase of generation is the most heterogeneous and therefore most informative. We also observe a systematic asymmetry between agent roles, with stronger alignment between confidence and quality for supportive reasoning than for adversarial critique. These results suggest that early decoding dynamics provide a lightweight and effective signal for estimating reasoning reliability in multi-agent LLM systems.

📄 PDF Abstract BibTeX arXiv:2606.10307

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding and Mitigating Premature Confidence for Better LLM Reasoning

2026-05-23 · Jingchu Gai, Guanning Zeng, Christina Baek, Chen Wu 외 arxiv

Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute. Improving reasoning quality directly would require pr…

Reinforcement Learning

Early Stopping for Large Reasoning Models via Confidence Dynamics

2026-04-06 · Parsa Hosseini, Sumit Nawathe, Mahdi Salmani, Meisam Razaviyayn 외 arxiv

Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. A key ch…

Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning

2025-10-09 · Aman Sharma, Paras Chopra arxiv

We introduce a simple, yet novel entropy-based framework to drive token efficiency in large language models during reasoning tasks. Our approach uses Shannon entropy from token-level logprobs as a confidence signal to en…

Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language Models

2025-09-28 · Zemin Huang, Yuhang Wang, Zhiyang Chen, Guo-Jun Qi arxiv

Mask-based Diffusion Language Models (DLMs) struggle to revise incorrect tokens: once a token is generated, it typically remains fixed. The key challenge is to identify potential errors in the inputs. In this paper, we p…

Reinforcement LearningText Generation

When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models

2026-06-29 · Zhe Dong, Fang Qin, Manish Shah arxiv

Reasoning models spend test-time compute unevenly across instances, and a growing family of early-exit rules -- confidence thresholds, entropy monitors, answer-stability checks, and learned stoppers -- promises to reclai…