paper-with-me

홈 › Papers

Trace Length is a Simple Uncertainty Signal in Reasoning Models

2025-10-12 · Siddartha Devic, Charlotte Peale, Arwen Bradley, Sinead Williamson, Preetum Nakkiran, Aravind Gollakota arxiv

Uncertainty quantification for LLMs is a key research direction towards addressing hallucination and other issues that limit their reliable deployment. In this work, we show that reasoning trace length is a simple and useful confidence estimator in large reasoning models. Through comprehensive experiments across multiple models, datasets, and prompts, we show that trace length performs in comparable but complementary ways to other zero-shot confidence estimators such as verbalized confidence. Our work reveals that reasoning post-training fundamentally alters the relationship between trace length and accuracy, going beyond prior work that had shown that post-training causes traces to grow longer in general (e.g., "overthinking"). We investigate the mechanisms behind trace length's performance as a confidence signal, observing that the effect remains even after adjusting for confounders such as problem difficulty and GRPO-induced length bias. We identify high-entropy or "forking" tokens as playing a key role in the mechanism. Our findings demonstrate that reasoning post-training enhances uncertainty quantification beyond verbal expressions, and establish trace length as a practical confidence measure for large reasoning models.

📄 PDF Abstract BibTeX arXiv:2510.10409

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio

2026-04-07 · Satwik Pandey, Suresh Raghu, Shashwat Pandey arxiv

Uncertainty estimation for reasoning language models remains difficult to deploy in practice: sampling-based methods are computationally expensive, while common single-pass proxies such as verbalized confidence or trace …

InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning

2026-03-18 · Chengwei Wei, Jung-jae Kim, Longyin Zhang, Shengkai Chen 외 arxiv

Large Language Models (LLMs) with extended reasoning capabilities often generate verbose and redundant reasoning traces, incurring unnecessary computational cost. While existing reinforcement learning approaches address …

Reinforcement Learning

SCOReD: Student-Aware CoT Optimization for Recommendation Distillation

2026-07-07 · Haz Sameen Shahgir, Yufei Li, Frank Shyu, Luke Simon 외 arxiv

Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-suited to this task. Large teachers approach the recommendation task with unusually…

Tracing Uncertainty in Language Model "Reasoning"

2026-05-08 · Nils Grünefeld, Bertram Højer, Philipp Mondorf, Barbara Plank 외 arxiv

Language model (LM) "reasoning", commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics underlying this process remain poorly understood. We study these dynami…

Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

2026-03-25 · Jeonghye Kim, Xufang Luo, Minbeom Kim, Sangmook Lee 외 arxiv

Self-distillation has emerged as an effective post-training paradigm for LLMs, often improving performance while shortening reasoning traces. However, in mathematical reasoning, we find that it can reduce response length…

Mathematical Reasoning