paper-with-me

홈 › Papers

Understanding and Mitigating Premature Confidence for Better LLM Reasoning

2026-05-23 · Jingchu Gai, Guanning Zeng, Christina Baek, Chen Wu, J. Zico Kolter, Andrej Risteski, Aditi Raghunathan arxiv

Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute. Improving reasoning quality directly would require process reward models, but the step-level annotations needed to train them are expensive and scarce. We find such a signal in how the model's confidence evolves during reasoning: premature confidence, the tendency to commit to an answer early and use the remaining tokens to rationalize it, strongly predicts flawed reasoning across tasks and model scales. We exploit this in progressive confidence shaping, a reinforcement learning objective that trains models to update their confidence as they reason rather than commit early -- rewarding gradual confidence growth and penalizing early commitment, with no external labels or reward models. The method improves accuracy and reasoning quality from 1.5B to 8B parameters across arithmetic (Countdown), math (DAPO, AIME), and science (ScienceQA): on Countdown, accuracy improves 3.2x (+42.0pp) and flawed reasoning drops 48pp; on AIME, Pass@64 improves 6.6pp. Consistent with this mechanism, the method also improves faithfulness: on a safety benchmark, our models more transparently surface misleading content in their reasoning traces rather than concealing it. Controlled experiments reveal that the problem and its remedy scale together: premature confidence grows with model size and task difficulty, and so do the gains from addressing it.

📄 PDF Abstract BibTeX arXiv:2605.24396

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning

2026-02-02 · Chu Zhao, Enneng Yang, Yuting Liu, Jianzhe Zhao 외 arxiv

Test-time reinforcement learning generates multiple candidate answers via repeated rollouts and performs online updates using pseudo-labels constructed by majority voting. To reduce overhead and improve exploration, prio…

Reinforcement LearningVisual Reasoning

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model

2025-06-23 · Xu Wan, Wei Wang, Wenyue Xu, Wotao Yin 외

Reinforcement Learning (RL)-based post-training has significantly advanced the complex reasoning capabilities of language models, fostering sophisticated self-reflection processes. However, this ``slow thinking'' paradig…

DiversityLanguage ModelingLanguage ModellingMathematical Reasoning+1

The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models

2026-05-27 · Dueun Kim, Albert No arxiv

Masked diffusion language models (MDMs) uniquely support any-order generation, with confidence-based decoding currently serving as the de facto standard inference policy. To optimize for this, recent training schemes att…

See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops

2025-08-25 · Zixuan Dong, Baoyun Peng, Yufei Wang, Lin Liu 외 arxiv

Human video comprehension demonstrates dynamic coordination between reasoning and visual attention, adaptively focusing on query-relevant details. However, current long-form video question answering systems employ rigid …

Video Question Answering

Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks

2025-12-01 · Jiannan Guan, Qiguang Chen, Libo Qin, Dengyun Peng 외 arxiv

Large Language Models (LLMs) excel in reasoning tasks requiring a single correct answer, but they perform poorly in multi-solution tasks that require generating comprehensive and diverse answers. We attribute this limita…