paper-with-me

홈 › Papers

PHF: Privileged Hidden Flow for On-Policy Self-Distillation

2026-06-28 · Yuhan Li, Mingxu Zhang, Dazhong Shen, Ying Sun arxiv

On-policy self-distillation (OPSD) trains a reasoning model on rollouts sampled from its own policy by matching a privileged teacher that also sees verified reference solutions. Existing OPSD objectives supervise only the output distribution, so privileged context affects training through a token-level divergence without directly supervising the internal computation that produced that distribution. We propose Privileged Hidden Flow (PHF), which additionally distills how a privileged teacher's hidden states move along the same rollout. Rather than forcing each student hidden vector to match the teacher vector at the same token position, PHF aligns token-to-token transition directions and trajectory geometry over selected generated positions. The all-layer recipe also includes an adjacent-layer relation computed from these same transitions, without pointwise hidden-state imitation. Under the same 100-step training schedule, PHF improves the Average@12 aggregate over our reproduced OPSD baseline on Qwen3-1.7B, 4B, and 8B, with observed gains of about +2.2, +1.5, and +1.7 points. The transport objective is exactly invariant to shared trajectory offsets; its local geometry term is also invariant to orthogonal transformations of transition directions. Ablations distinguish the fixed PHF recipe from pointwise hidden-state matching, single-channel transition losses, and layer-subset choices, supporting PHF as a compact hidden-flow extension to OPSD.

📄 PDF Abstract BibTeX arXiv:2606.29340

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rethinking On-Policy Self-Distillation for Thinking Models

2026-07-06 · Simran Kaur, Narutatsu Ri, Yinghui He, Liam Fowl 외 arxiv

Self-distillation is a promising recipe for self-improvement in language models. In this setting, a model can serve as its own teacher when given privileged information, such as a solution to a math problem. This seems e…

Latent On-Policy Self-Distillation

2026-08-13 · Guibin Zhang, Jiayang Lyu, Ran Sun, Xinlei Yu 외 arxiv

Enabling agents to learn from experience and internalize it into their policy has become a central problem in self-evolving AI. On-policy self-distillation (OPSD) offers an effective pathway by using a privileged self-te…

Code Generation

HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation

2026-03-25 · Ken Ding arxiv

Large language models trained with reinforcement learning (RL) for mathematical reasoning face a fundamental challenge: on problems the model cannot solve at all - "cliff" prompts - the RL gradient vanishes entirely, pre…

Reinforcement LearningMathematical Reasoning

Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning

2026-05-21 · Hongbin Zhang, Chaozheng Wang, Kehai Chen, Youcheng Pan 외 arxiv

On-policy self-distillation (OPSD) is an emerging LLM post-training paradigm in which the model serves as its own teacher: conditioned on privileged information such as a reference trace or hint, the same policy provides…

Mathematical Reasoning

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

2026-01-26 · Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang 외 arxiv

Knowledge distillation improves large language model (LLM) reasoning by compressing the knowledge of a teacher LLM to train smaller LLMs. On-policy distillation advances this approach by having the student sample its own…

Knowledge DistillationMathematical ReasoningReinforcement Learning