paper-with-me

홈 › Papers

Lexical Hints of Accuracy in LLM Reasoning Chains

2025-08-19 · Arne Vanhoyweghen, Brecht Verbeken, Andres Algaba, Vincent Ginis arxiv

Fine-tuning Large Language Models (LLMs) with reinforcement learning to produce an explicit Chain-of-Thought (CoT) before answering produces models that consistently raise overall performance on code, math, and general-knowledge benchmarks. However, on benchmarks where LLMs currently achieve low accuracy, such as Humanity's Last Exam (HLE), they often report high self-confidence, reflecting poor calibration. Here, we test whether measurable properties of the CoT provide reliable signals of an LLM's internal confidence in its answers. We analyze three feature classes: (i) CoT length, (ii) intra-CoT sentiment volatility, and (iii) lexicographic hints, including hedging words. Using DeepSeek-R1 and Claude 3.7 Sonnet on both Humanity's Last Exam (HLE), a frontier benchmark with very low accuracy, and Omni-MATH, a saturated benchmark of moderate difficulty, we find that lexical markers of uncertainty (e.g., $\textit{guess}$, $\textit{stuck}$, $\textit{hard}$) in the CoT are the strongest indicators of an incorrect response, while shifts in the CoT sentiment provide a weaker but complementary signal. CoT length is informative only on Omni-MATH, where accuracy is already high ($\approx 70\%$), and carries no signal on the harder HLE ($\approx 9\%$), indicating that CoT length predicts correctness only in the intermediate-difficulty benchmarks, i.e., inside the model's demonstrated capability, but still below saturation. Finally, we find that uncertainty indicators in the CoT are consistently more salient than high-confidence markers, making errors easier to predict than correct responses. Our findings support a lightweight post-hoc calibration signal that complements unreliable self-reported probabilities and supports safer deployment of LLMs.

📄 PDF Abstract BibTeX arXiv:2508.15842

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

2025-07-03 · Kaiyi Zhang, Ang Lv, Jinpeng Li, Yongbo Wang 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) is a promising approach for improving the complex reasoning abilities of large language models (LLMs). However, current RLVR methods face two significant challenges: …

Reinforcement Learning

ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better

2025-11-21 · Yuan Zhang, Ming Lu, Junwen Pan, Tao Huang 외 arxiv

Recent advances in multimodal reasoning models have demonstrated impressive capabilities across text and vision. However, even leading models exhibit redundant self-reflection when generating lengthy reasoning chains. Wh…

Multimodal Reasoning

Unleashing the Creative Mind: Language Model As Hierarchical Policy For Improved Exploration on Challenging Problem Solving

2023-11-01 · Zhan Ling, Yunhao Fang, Xuanlin Li, Tongzhou Mu 외

Large Language Models (LLMs) have achieved tremendous progress, yet they still often struggle with challenging reasoning problems. Current approaches address this challenge by sampling or searching detailed and low-level…

In-Context LearningLanguage ModelingLanguage ModellingMath

Length Penalties Make Chain-of-Thought Less Monitorable

2026-07-08 · Bryce Little hf

Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints f…

Reinforcement Learning

HintMR: Eliciting Stronger Mathematical Reasoning in Small Language Models

2026-04-14 · Jawad Hossain, Xiangyu Guo, Jiawei Zhou, Chong Liu arxiv

Small language models (SLMs) often struggle with complex mathematical reasoning due to limited capacity to maintain long chains of intermediate steps and to recover from early errors. We address this challenge by introdu…

Mathematical Reasoning