paper-with-me

홈 › Papers

Scaling Reasoning Efficiently via Relaxed On-Policy Distillation

2026-03-11 · Jongwoo Ko, Sara Abdali, Young Jin Kim, Tianyi Chen, Pashmina Cameron arxiv

On-policy distillation is pivotal for transferring reasoning capabilities to capacity-constrained models, yet remains prone to instability and negative transfer. We show that on-policy distillation can be interpreted, both theoretically and empirically, as a form of policy optimization, where the teacher-student log-likelihood ratio acts as a token reward. From this insight, we introduce REOPOLD (Relaxed On-Policy Distillation) a framework that stabilizes optimization by relaxing the strict imitation constraints of standard on-policy distillation. Specifically, REOPOLD temperately and selectively leverages rewards from the teacher through mixture-based reward clipping, entropy-based token-level dynamic sampling, and a unified exploration-to-refinement training strategy. Empirically, REOPOLD surpasses its baselines with superior sample efficiency during training and enhanced test-time scaling at inference, across mathematical, visual, and agentic tool-use reasoning tasks. Specifically, REOPOLD outperforms recent RL approaches achieving 6.7~12x greater sample efficiency and enables a 7B student to match a 32B teacher in visual reasoning with a ~3.32x inference speedup.

📄 PDF Abstract BibTeX arXiv:2603.11137

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

2025-08-13 · Xiaojun Wu, Xiaoguang Jiang, Huiyang Li, Jucai Zhai 외 arxiv

Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solving. Recent methods have improved reasoning through expanded corpus and multista…

Reinforcement LearningMathematical ReasoningCode Generation

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

2026-07-30 · Jiawei Xu, Minghui Liu, Juzheng Zhang, Tom Goldstein 외 arxiv

On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a st…

Mathematical ReasoningReinforcement Learning

On the Geometry of On-Policy Distillation

2026-06-05 · Zhennan Shen, Yanshu Li, Qingyu Yin, Chak Tou Leong 외 arxiv

On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly understood. We characterize the trajectory of OPD updates in parameter space and compar…

Reinforcement Learning

Hán Dān Xué Bù (Mimicry) or Qīng Chū Yú Lán (Mastery)? A Cognitive Perspective on Reasoning Distillation in Large Language Models

2026-01-08 · Yueqing Hu, Xinyang Peng, Shuting Peng, Hanqi Wang 외 arxiv

Recent Large Reasoning Models trained via reinforcement learning exhibit a "natural" alignment with human cognitive costs. However, we show that the prevailing paradigm of reasoning distillation -- training student model…

Reinforcement Learning

Crosslingual On-Policy Self-Distillation for Multilingual Reasoning

2026-05-10 · Yihong Liu, Raoyuan Zhao, Michael A. Hedderich, Hinrich Schütze arxiv

Large language models (LLMs) have achieved remarkable progress in mathematical reasoning, but this ability is not equally accessible across languages. Especially low-resource languages exhibit much lower reasoning perfor…

Reinforcement LearningMathematical Reasoning