paper-with-me

홈 › Papers

Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning

2026-03-11 · Zichao Li, Jie Lou, Fangchen Dong, Zhiyuan Fan, Mengjie Ren, Hongyu Lin, Xianpei Han, Debing Zhang, Le Sun, Yaojie Lu, Xing Yu arxiv

Reinforcement learning significantly enhances LLM capabilities but suffers from a critical issue: length inflation, where models adopt verbosity or inefficient reasoning to maximize rewards. Prior approaches struggle to address this challenge in a general and lossless manner, primarily because additive penalties introduce a compensatory effect that creates optimization shortcuts, while heuristic gating strategies lack generality beyond binary feedback. To bridge this gap, we present Group Relative Reward Rescaling (GR$^3$), which reframes length control as a multiplicative rescaling paradigm, effectively establishing a generalized, continuous, and reward-dependent gating mechanism. To further ensure lossless optimization, we incorporate group-relative regularization and advantage-aware calibration, which dynamically adapt length budgets to instance difficulty and preserve the advantage signal of high-quality trajectories. Empirically, across both RLHF and RLVR settings, GR$^3$~maintains training dynamics and downstream performance comparable to standard GRPO while significantly mitigating length inflation, outperforming state-of-the-art length-regularized baselines.

📄 PDF Abstract BibTeX arXiv:2603.10535

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models

2026-06-24 · Xinyu Lian, Walid Krichene, Beichen Huang, Masahiro Tanaka 외 arxiv

Quantization is widely used to reduce the inference cost of large language models, but its effect on reasoning models is not fully captured by final-answer accuracy or per-token latency. We show that low-bit post-trainin…

Mathematical ReasoningQuestion AnsweringCode Generation

Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation

2026-05-16 · Anhao Zhao, Haoran Xin, Yingqi Fan, Junlong Tong 외 arxiv

Knowledge distillation is central to LLM post-training, yet its design space remains poorly understood, especially alongside reinforcement learning (RL). We show that the prevailing paradigms, off-policy distillation and…

Knowledge DistillationReinforcement LearningOffline RL

Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

2025-08-13 · Vaishnavi Shrivastava, Ahmed Awadallah, Vidhisha Balachandran, Shivam Garg 외 arxiv

Large language models trained with reinforcement learning with verifiable rewards tend to trade accuracy for length--inflating response lengths to achieve gains in accuracy. While longer answers may be warranted for hard…

Computational EfficiencyReinforcement Learning

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

2026-09-17 · Yuxiao Yang, Tianrun Yu, Shangzhe Li, Kaixiang Zhao 외 hf

We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post…

EST-PRM: Stress-Testing Process Reward Models Before They Become Load-Bearing

2026-05-30 · Ibne Farabi Shihab, Fariya Afrin, Sanjeda Akter, Anuj Sharma arxiv

Process reward models (PRMs) are widely used in language-model training with dense step-level supervision. They assume PRM scores are stable proxies for step correctness under label-preserving transformations. These tran…