paper-with-me

홈 › Papers

Distribution Matching Distillation Meets Reinforcement Learning

2025-11-17 · Dengyang Jiang, Dongyang Liu, Zanyi Wang, Qilong Wu, Liuzhuozheng Li, Hengzhuang Li, Xin Jin, David Liu, Changsheng Lu, Zhen Li, Bo Zhang, Mengmeng Wang, Steven Hoi, Peng Gao, Harry Yang arxiv

Distribution Matching Distillation (DMD) facilitates efficient inference by distilling multi-step diffusion models into few-step variants. Concurrently, Reinforcement Learning (RL) has emerged as a vital tool for aligning generative models with human preferences. While both represent critical post-training stages for large-scale diffusion models, existing studies typically treat them as independent, sequential processes, leaving a systematic framework for their unification largely unexplored. In this work, we demonstrate that jointly optimizing these two objectives yields mutual benefits: RL enables more preference-aware and controllable distillation rather than uniformly compressing the full data distribution, while DMD serves as an effective regularizer to mitigate reward hacking during RL training. Building on these insights, we propose DMDR, a unified framework that incorporates Reward-Tilted Distribution Matching optimization alongside two dynamic distillation training strategies in the initial stage, followed by the joint DMD and RL optimization in the second stage. Extensive experiments demonstrate that DMDR achieves state-of-the-art visual quality and prompt adherence among few-step generation methods, even surpassing the performance of its multi-step teacher model.

📄 PDF Abstract BibTeX arXiv:2511.13649

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

AdvDMD: Adversarial Reward Meets DMD For High-Quality Few-Step Generation

2026-04-29 · Xu Wang, Zexian Li, Litong Gong, Tiezheng Ge 외 arxiv

Diffusion models offer superior generation quality at the expense of extensive sampling steps. Distillation methods, with Distribution Matching Distillation (DMD) as a popular example, can mitigate this issue, but perfor…

Reinforcement Learning

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching

2026-05-25 · Yushi Huang, Xiangxin Zhou, Ruoyu Wang, Chi Zhang 외 arxiv

Recent advances in few-step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains challenging. We propose Reward-Tilted Distribution Matching Distillatio…

Text-to-Image GenerationReinforcement Learning

Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning

2026-04-21 · Linwei Dong, Ruoyu Guo, Ge Bai, Zehuan Yuan 외 arxiv

Diffusion distillation, exemplified by Distribution Matching Distillation (DMD), has shown great promise in few-step generation but often sacrifices quality for sampling speed. While integrating Reinforcement Learning (R…

Reinforcement Learning

Multi-Source Domain Adaptation meets Dataset Distillation through Dataset Dictionary Learning

2023-09-14 · Eduardo Fernandes Montesuma, Fred Ngolè Mboula, Antoine Souloumiac

In this paper, we consider the intersection of two problems in machine learning: Multi-Source Domain Adaptation (MSDA) and Dataset Distillation (DD). On the one hand, the first considers adapting multiple heterogeneous l…

Dataset DistillationDictionary LearningDomain Adaptation

Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

2026-06-21 · Songtao Tian, Guhan Chen, Bohan Li, Jingyi Ma 외 arxiv

Consistency distillation has significantly accelerated the inference of diffusion models. In this work, we reveal an intriguing asymmetry: while Logit-Normal sampling priors are highly efficacious for standard iterative …

Reinforcement Learning