paper-with-me

홈 › Papers

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

2026-08-17 · Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling, Jiahang Xie, Zhiyong Zheng, Yujia Wang, Hao Liu, Feng Xiao, Lu Liu, Yanlong Du, Zifeng Cheng, Ziwei Jiang, Qing Gu arxiv

On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy distillation with substantially less data. However, standard token-level OPD can provide only fragmented corrections along an erroneous student trajectory and cannot unfold a complete and correct repair path. Motivated by this limitation, we propose \emph{Step-Level On-Policy Distillation} (SOPD), which combines the long-horizon correction of supervised fine-tuning (SFT) with the on-policy advantage of OPD to provide step-level supervision over complete student-generated trajectories. We show that, at different limits of step length, SOPD reduces to SFT or approximates OPD. Compared with SFT, the teacher responses in SOPD are conditioned on student trajectories and therefore align more closely with student-visited states; compared with OPD, SOPD provides longer-horizon corrections rather than fragmented token-level guidance. Across both reasoning and agent tasks, SOPD substantially outperforms conventional SFT and OPD. For example, on ALFWorld, SOPD improves the average success rate by 13.4 points over Vanilla OPD. We hope this work offers a new perspective for future research on distillation methods.

📄 PDF Abstract BibTeX arXiv:2608.16333

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SOD: Step-wise On-policy Distillation for Small Language Model Agents

2026-05-08 · Qiyong Zhong, Mao Zheng, Mingyang Song, Xin Lin 외 arxiv

Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and limited model capacity. While reinforcement learning methods like group relative pol…

Reinforcement Learning

MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

2026-05-02 · Jianze Wang, Ying Liu, Jinlong Chen, Xuchun Hu 외 arxiv

On-policy distillation (OPD) trains a student on its own trajectories under token-level teacher supervision, but existing methods are capped by a single-teacher capability ceiling: when the teacher errs, the student inhe…

Code Generation

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

2026-09-04 · Yang Li, Semih Yavuz, Shafiq Joty arxiv

On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectiveness is bottlenecked by teacher quality: external teachers suffer from distribution mismatch, while se…

Mathematical ReasoningCode Generation

Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization

2026-06-05 · Shizhe Xiang, Ke An, Wenlong Yu, Yue Liu 외 arxiv

Recent post-training methods, particularly Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced the reasoning ability of Large Vision-Language Models (LVLMs). However, the sparse nature of v…

Reinforcement LearningMultimodal Reasoning

OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

2026-06-25 · Shuo Yang, Jinyang Wu, Zhengxi Lu, Yuhao Shen 외 arxiv

Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse trajectory-level rewards provide little guidance on which intermediate decisions should be reinforced or su…

Reinforcement Learning