paper-with-me

홈 › Papers

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

2026-04-09 · Feng Luo, Yu-Neng Chuang, Guanchu Wang, Zicheng Xu, Xiaotian Han, Tianyi Zhang, Vladimir Braverman arxiv

On-policy distillation (OPD) trains student models under their own induced distribution while leveraging supervision from stronger teachers. We identify a failure mode of OPD: as training progresses, on-policy rollouts can undergo abrupt length inflation, causing truncated trajectories to dominate the training data. This truncation collapse coincides with abrupt repetition saturation and induces biased gradient signals, leading to severe training instability and sharp degradation in validation performance. We attribute this problem to the interaction between student-induced data collection and the distillation objective, which implicitly favors long and repetitive rollouts. To address this issue, we propose StableOPD, a stabilized OPD framework that combines a reference-based divergence constraint with rollout mixture distillation. These together mitigate repetition-induced length inflation and further stabilize OPD training. Across multiple math reasoning datasets, our approach prevents truncation collapse, stabilizes training dynamics, and improves performance by 7.2% on average.

📄 PDF Abstract BibTeX arXiv:2604.08527

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bounded Rationality in Central Bank Communication

2024-11-06 · Wonseong Kim, Choong Lyol Lee

This study explores the influence of FOMC sentiment on market expectations, focusing on cognitive differences between experts and non-experts. Using sentiment analysis of FOMC minutes, we integrate these insights into a …

Sentiment Analysis

Who's at Risk? Effects of Inflation on Unemployment Risk

2025-05-09 · Hie Joo Ahn, Lam Nguyen

We empirically investigate the distributional effects of inflation on workers' unemployment tail risks using instrumental variable quantile regression. We find that supply-driven inflation disproportionately raises unemp…

quantile regression

Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning

2026-03-11 · Zichao Li, Jie Lou, Fangchen Dong, Zhiyuan Fan 외 arxiv

Reinforcement learning significantly enhances LLM capabilities but suffers from a critical issue: length inflation, where models adopt verbosity or inefficient reasoning to maximize rewards. Prior approaches struggle to …

Reinforcement Learning

Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models

2026-06-24 · Xinyu Lian, Walid Krichene, Beichen Huang, Masahiro Tanaka 외 arxiv

Quantization is widely used to reduce the inference cost of large language models, but its effect on reasoning models is not fully captured by final-answer accuracy or per-token latency. We show that low-bit post-trainin…

Mathematical ReasoningQuestion AnsweringCode Generation

Demystifying Workload Imbalances in Large Transformer Model Training over Variable-length Sequences

2024-12-10 · Haoyang Li, Fangcheng Fu, Sheng Lin, Hao Ge 외

To optimize large Transformer model training, efficient parallel computing and advanced data management are essential. However, current methods often assume a stable and uniform training workload, neglecting imbalances i…

Management