paper-with-me

홈 › Papers

Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning

2026-02-27 · Xintong Li, Sha Li, Rongmei Lin, Hongye Jin, Linwei Li, Hejie Cui, Sarah Zhang, Chia-Yuan Chang, Kewei Cheng, Besnik Fetahu, Priyanka Nigam, Jingbo Shang, Bing Yin arxiv

Large reasoning models improve with more test-time computation, but often overthink, producing unnecessarily long chains-of-thought that raise cost without improving accuracy. Prior reinforcement learning approaches typically rely on a single outcome reward with trajectory-level length penalties, which cannot distinguish essential from redundant reasoning steps and therefore yield blunt compression. Although recent work incorporates step-level signals, such as offline pruning, supervised data construction, or verifier-based intermediate rewards, reasoning length is rarely treated as an explicit step-level optimization objective during RL. We propose Step-wise Adaptive Penalization (SWAP), a fine-grained framework that allocates length reduction across steps based on intrinsic contribution. We estimate step importance from the model's on-policy log-probability improvement toward the correct answer, then treat excess length as a penalty mass redistributed to penalize low-importance steps more heavily while preserving high-importance reasoning. We optimize with a unified outcome-process advantage within group-relative policy optimization. Extensive experiments demonstrate that SWAP reduces reasoning length by 64.3% on average while improving accuracy by 5.7% relative to the base model.

📄 PDF Abstract BibTeX arXiv:2603.00296

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization

2025-06-12 · Zhensheng Jin, Xinze Li, Yifan Ji, Chunyi Peng 외

Recent advances in Chain-of-Thought (CoT) prompting have substantially improved the reasoning capabilities of Large Language Models (LLMs). However, these methods often suffer from overthinking, leading to unnecessarily …

Math

SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

2026-03-09 · Chenzhi Hu, Qinzhe Hu, Yuhang Xu, Junyi Chen 외 arxiv

Large reasoning models (LRMs) like OpenAI o1 and DeepSeek-R1 achieve high accuracy on complex tasks by adopting long chain-of-thought (CoT) reasoning paths. However, the inherent verbosity of these processes frequently r…

HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression

2026-06-01 · Minghui Zheng, Hongxu Chen, Huimin Ren, Hongsheng Xin 외 arxiv

Large language models achieve remarkable performance via extended chain-of-thought (CoT) reasoning, yet this lengthy process incurs substantial inference overhead. Existing CoT compression methods struggle with inflexibl…

Reinforcement Learning

A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware Demonstration

2024-10-21 · Yingqian Cui, Pengfei He, Xianfeng Tang, Qi He 외

Few-shot Chain-of-Thought (CoT) prompting has demonstrated strong performance in improving the reasoning capabilities of large language models (LLMs). While theoretical investigations have been conducted to understand Co…

In-Context Learning

VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

2025-04-10 · Yukun Qi, Yiming Zhao, Yu Zeng, Xikun Bao 외

The advancement of Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs) and large vision-language models (LVLMs). However, a rigorous evaluation framework for video…