paper-with-me

홈 › Papers

Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation

2026-06-16 · Siyi Chen, Shaowei Liu, Yixuan Jia, Zian Wang, Huan Ling, Qing Qu, Jun Gao arxiv

Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matching Distillation (DMD) and its successor DMD2 achieved strong generation quality and fast convergence. However, due to the nature of the reverse Kullback--Leibler (KL) objective, these methods exhibit two persistent failure modes: a substantial drop in sample diversity, and visibly over-saturated outputs that deviate from real-video appearance. In this work, we propose Data-Forcing Distillation (DFD), a simple post-training framework that restores diversity and fidelity in DMD with only a single-line of code change. At its core is the teacher score discrepancy to guide the student toward the real-data distribution, pulling it to missing modes (mitigating mode collapse) and away from problematic modes absent in real data (avoiding over-saturation). We provide an in-depth theoretical analysis of our framework and validate our approach on text-to-video, image-to-video, and autoregressive video generation. With only 100--300 steps of finetuning, DFD effectively restores diversity and fidelity on both Wan2.1-1.3B and Cosmos-Predict2.5-2B model, resolving the over-saturation artifacts with significantly better video dynamics and appearance, and even outperforms the teacher model.

📄 PDF Abstract BibTeX arXiv:2606.18478

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Restoring Initial Noise Sensitivity in Text-to-Image Distillation via Geometric Alignment

2026-06-01 · Huayang Huang, Ruoyu Wang, Jinhui Zhao, Wei Deng 외 arxiv

Generative distillation significantly accelerates text-to-image (T2I) generation by compressing multi-step trajectories into few-step student models while preserving perceptual quality. However, existing methods primaril…

Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation

2026-03-31 · Hoang-Chau Luong, Dat Ba Tran, Lingwei Chen arxiv

Reverse Kullback-Leibler (RKL) divergence has recently emerged as the preferred objective for large language model (LLM) distillation, consistently outperforming forward KL (FKL), particularly in regimes with large vocab…

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation

2026-07-08 · Jiabin Shen, Guang Chen, Chengjun Mao arxiv

Top-$K$ teacher logits make on-policy distillation tractable, but retained teacher mass does not certify student-relative gradient fidelity. We study routed, two-teacher tool-use distillation. In a frozen Qwen3.5-9B audi…

Knowledge Distillation

Exploiting Knowledge Distillation for Few-Shot Image Generation

2021-09-29 · Xingzhong Hou, Boxiao Liu, Fang Wan, Haihang You

Few-shot image generation, which trains generative models on limited examples, is of practical importance. The existing pipeline is first pretraining a source model (which contains a generator and a discriminator) on a l…

DiversityImage GenerationKnowledge DistillationRelation

Multi-fidelity Neural Architecture Search with Knowledge Distillation

2020-06-15 · Ilya Trofimov, Nikita Klyuchnikov, Mikhail Salnikov, Alexander Filippov 외

Neural architecture search (NAS) targets at finding the optimal architecture of a neural network for a problem or a family of problems. Evaluations of neural architectures are very time-consuming. One of the possible way…

Knowledge DistillationNeural Architecture Search