paper-with-me

홈 › Papers

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level

2026-05-07 · Nan Jia, Haojin Yang, Xing Ma, Jiesong Lian, Shuailiang Zhang, Weipeng Zhang, Ke Zeng, Xunliang Cai, Zequn Sun arxiv

On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher feedback and often outperforms off-policy distillation and standard reinforcement learning. However, we find that its standard advantage weighted policy gradient suffers from three structural weaknesses, including high variance updates, vanishing gradients in zero-advantage regions, and exploration bottlenecks when corrective signals are insufficient. We therefore propose Asymmetric On-Policy Distillation (AOPD), which replaces ineffective negative reinforcement with localized divergence minimization in non-positive advantage regions while preserving positive reinforcement learning. Experiments on mathematical reasoning benchmarks show that AOPD consistently outperforms standard OPD, with average gains of 4.09 / 8.34 under strong / weak initialization, respectively. AOPD also maintains higher policy entropy during training and better capability retention during sequential tool-use adaptation.

📄 PDF Abstract BibTeX arXiv:2605.06387

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical Reasoning

Similar Papers 제목 키워드 기반

VEPO: Variable Entropy Policy Optimization for Low-Resource Language Foundation Models

2026-03-19 · Chonghan Liu, Yimin Du, Qi An, Xin He 외 arxiv

Large language models frequently exhibit suboptimal performance on low resource languages, primarily due to inefficient subword segmentation and systemic training data imbalances. In this paper, we propose Variable Entro…

Reinforcement Learning

Reinforcement-aware Knowledge Distillation for LLM Reasoning

2026-02-26 · Zhaoyang Zhang, Shuli Jiang, Yantao Shen, Yuting Zhang 외 arxiv

Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs), but the high inference cost of such models motivates distillation into smaller st…

Knowledge DistillationReinforcement Learning

GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies

2026-03-15 · He Zhang, Ying Sun, Hui Xiong arxiv

Flow-matching policies hold great promise for reinforcement learning (RL) by capturing complex, multi-modal action distributions. However, their practical application is often hindered by prohibitive inference latency an…

Reinforcement LearningContinuous Control

AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment

2026-05-18 · Zhenlin Wei, Pu Jian, Yingzhuo Deng, Xiaohan Wang 외 arxiv

The alignment of Large Language Models (LLMs) for complex reasoning heavily relies on Reinforcement Learning with Verifiable Rewards (RLVR). However, standard algorithms like GRPO apply sequence-level rewards uniformly t…

Reinforcement Learning

Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency

2025-11-12 · Riling Wei, Kelu Yao, Chuanguang Yang, Jin Wang 외 arxiv

Cross-modal Knowledge Distillation has demonstrated promising performance on paired modalities with strong semantic connections, referred to as Symmetric Cross-modal Knowledge Distillation (SCKD). However, implementing S…

Self-Supervised LearningKnowledge DistillationScene Classification