paper-with-me

홈 › Papers

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

2025-06-05 · Violet Xiang, Chase Blagden, Rafael Rafailov, Nathan Lile, Sang Truong, Chelsea Finn, Nick Haber

Large reasoning models (LRMs) achieve higher performance on challenging reasoning tasks by generating more tokens at inference time, but this verbosity often wastes computation on easy problems. Existing solutions, including supervised finetuning on shorter traces, user-controlled budgets, or RL with uniform penalties, either require data curation, manual configuration, or treat all problems alike regardless of difficulty. We introduce Adaptive Length Penalty (ALP), a reinforcement learning objective tailoring generation length to per-prompt solve rate. During training, ALP monitors each prompt's online solve rate through multiple rollouts and adds a differentiable penalty whose magnitude scales inversely with that rate, so confident (easy) prompts incur a high cost for extra tokens while hard prompts remain unhindered. Posttraining DeepScaleR-1.5B with ALP cuts average token usage by 50\% without significantly dropping performance. Relative to fixed-budget and uniform penalty baselines, ALP redistributes its reduced budget more intelligently by cutting compute on easy prompts and reallocating saved tokens to difficult ones, delivering higher accuracy on the hardest problems with higher cost.

📄 PDF Abstract BibTeX arXiv:2506.05256

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models

2025-03-06 · Yi Shen, Jian Zhang, Jieyun Huang, Shuming Shi 외

Recent advancements in slow-thinking reasoning models have shown exceptional performance in complex reasoning tasks. However, these models often exhibit overthinking-generating redundant reasoning steps for simple proble…

Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking

2025-09-27 · Jinyi Han, Ying Huang, Ying Liao, Zishang Jiang 외 arxiv

Large Reasoning Models (LRMs) have achieved impressive performance on challenging tasks, yet their deep reasoning often incurs substantial computational costs. To achieve efficient reasoning, existing reinforcement learn…

Reinforcement Learning

Stop When Enough: Adaptive Early-Stopping for Chain-of-Thought Reasoning

2025-10-11 · Renliang Sun, Wei Cheng, Dawei Li, Haifeng Chen 외 arxiv

Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. However, excessive or redundant reasoning -- so-called overthinki…

Output Length Effect on DeepSeek-R1's Safety in Forced Thinking

2025-03-02 · Xuying Li, Zhuo Li, Yuji Kosuga, Victor Bian

Large Language Models (LLMs) have demonstrated strong reasoning capabilities, but their safety under adversarial conditions remains a challenge. This study examines the impact of output length on the robustness of DeepSe…

Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression

2025-10-02 · Joykirat Singh, Justin Chih-Yao Chen, Archiki Prasad, Elias Stengel-Eskin 외 arxiv

Recent thinking models solve complex reasoning tasks by scaling test-time compute, but this scaling must be allocated in line with task difficulty. On one hand, short reasoning (underthinking) leads to errors on harder p…