paper-with-me

홈 › Papers

D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models

2026-05-06 · Dengyang Jiang, Xin Jin, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Ruoyi Du, Xiangpeng Yang, Qilong Wu, Zhen Li, Peng Gao, Harry Yang, Steven Hoi arxiv

The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and FLUX.2-klein). However, these models present significant challenges for direct continuous supervised fine-tuning. For example, applying the commonly used fine-tuning technique would compromise their inherent few-step inference capability. To address this, we propose D-OPSD, a novel training paradigm for step-distilled diffusion models that enables on-policy learning during supervised fine-tuning. We first find that the modern diffusion models, where the LLM/VLM serves as the encoder, can inherit its encoder's in-context capabilities. This enables us to formulate the training as an on-policy self-distillation process. Specifically, during training, we make the model act as both the teacher and the student with different contexts, where the student is conditioned only on the text feature, while the teacher is conditioned on the multimodal feature of both the text prompt and the target image. Training minimizes the two predicted distributions over the student's own roll-outs. By optimizing on the model's own trajectory and under its own supervision, D-OPSD enables the model to learn new concepts, styles, etc., without sacrificing the original few-step capacity.

📄 PDF Abstract BibTeX arXiv:2605.05204

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

On-Policy Self-Distillation without Any Supervision

2026-08-06 · Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian 외 arxiv

On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, …

Mathematical Reasoning

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

2026-07-30 · Jiawei Xu, Minghui Liu, Juzheng Zhang, Tom Goldstein 외 arxiv

On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a st…

Mathematical ReasoningReinforcement Learning

dOPSD: On-Policy Self-Distillation for Diffusion Language Models

2026-07-05 · Phuong Tuan Dat, Qi Li, Xinchao Wang hf

Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains dif…

Reinforcement LearningCode Generation

RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer

2026-08-06 · Xinye Wang, Junxiao Liu, Shujian Huang arxiv

Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-policy self-distillation (OPSD) and its variants have emerged as a promis…

Mathematical ReasoningCross-Lingual Transfer

DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation

2026-08-26 · Yutong Chen, Guangfu Guo, Zhichao Xu, Kunpeng Liu arxiv

On-policy self-distillation (OPSD) uses a privileged copy of the student model to provide dense supervision without an external teacher. OPSD keeps this privileged teacher fixed, even though the student distribution and …