paper-with-me

홈 › Papers

$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control

2026-05-18 · Xianwei Chen, Shimin Zhang, Jibin Wu arxiv

Scaling on-policy distillation (OPD) for large language models (LLMs) confronts a fundamental tension: asynchronous execution is necessary for system efficiency, but structurally deviates from the ideal on-policy objective. To address this challenge, we theoretically decompose the objective discrepancy into rollout drift and supervision drift, capturing staleness in student rollout and teacher context, respectively. Building on this, we introduce a sample-level freshness score that quantifies the reliability of a buffered sample with respect to the on-policy objective. Guided by this signal, we further propose f-OPD, a novel framework that adaptively regulates stale-sample influence and constrains policy drift accumulated under asynchronous training. Across reasoning, tool-use, and coding-agent tasks of increasing interaction horizon, f-OPD consistently achieves task performance comparable to synchronous optimization while largely retaining the throughput advantages of asynchronous execution. Our results establish the first recipe for achieving a performance-efficiency trade-off in OPD, paving the way for long-horizon agentic post-training at scale.

📄 PDF Abstract BibTeX arXiv:2605.17862

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adversarial Dual On-Policy Distillation from Expressive Teacher

2026-05-26 · Zhenglin Wan, Jingxuan Wu, Xingrui Yu, Chubin Zhang 외 arxiv

Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradigm by modeling multi-modal expert actions. Yet these methods remain …

Robot Navigation

Revisiting LQR Control from the Perspective of Receding-Horizon Policy Gradient

2023-02-25 · Xiangyuan Zhang, Tamer Başar

We revisit in this paper the discrete-time linear quadratic regulator (LQR) problem from the perspective of receding-horizon policy gradient (RHPG), a newly developed model-free learning framework for control application…

Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning

2026-08-17 · Changhui Sun, Lanbo Liu, Hang Lei, Tong Ling 외 arxiv

On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach has achieved strong empirical gains and can often surpass conventional off-policy d…

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

2026-07-07 · Yuhang Zhou, Kai Zheng, Haoling Li, Dengyun Peng 외 arxiv

On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon …

On receding-horizon approximation in time-varying optimal control

2023-05-10 · Jintao Sun, Michael Cantoni

The closed-loop stability and infinite-horizon performance of receding-horizon approximations are studied for non-stationary linear-quadratic regulator (LQR) problems. The approach is based on a lifted reformulation of t…