paper-with-me

Papers

Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning

2025-11-25 · Guanjie Chen, Shirui Huang, Kai Liu, Jianchen Zhu, Xiaoye Qu, Peng Chen, Yu Cheng, Yifu Sun arxiv

Diffusion Models have emerged as a leading class of generative models, yet their iterative sampling process remains computationally expensive. Timestep distillation is a promising technique to accelerate generation, but it often requires extensive training and leads to image quality degradation. Furthermore, fine-tuning these distilled models for specific objectives, such as aesthetic appeal or user preference, using Reinforcement Learning (RL) is notoriously unstable and easily falls into reward hacking. In this work, we introduce Flash-DMD, a novel framework that enables fast convergence with distillation and joint RL-based refinement. Specifically, we first propose an efficient timestep-aware distillation strategy that significantly reduces training cost with enhanced realism, outperforming DMD2 with only $2.1\%$ its training cost. Second, we introduce a joint training scheme where the model is fine-tuned with an RL objective while the timestep distillation training continues simultaneously. We demonstrate that the stable, well-defined loss from the ongoing distillation acts as a powerful regularizer, effectively stabilizing the RL training process and preventing policy collapse. Extensive experiments on score-based and flow matching models show that our proposed Flash-DMD not only converges significantly faster but also achieves state-of-the-art generation quality in the few-step sampling regime, outperforming existing methods in visual quality, human preference, and text-image alignment metrics. Our work presents an effective paradigm for training efficient, high-fidelity, and stable generative models. Codes are coming soon.

📄 PDF Abstract BibTeX arXiv:2511.20549

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage Generation

Similar Papers 제목 키워드 기반

FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation

2024-10-16 · Huadai Liu, Jialei Wang, Rongjie Huang, Yang Liu 외

Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While…

Audio GenerationGPU

Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image Generation

2024-06-04 · Clément Chadebec, Onur Tasar, Eyal Benaroche, Benjamin Aubin

In this paper, we propose an efficient, fast, and versatile distillation method to accelerate the generation of pre-trained diffusion models: Flash Diffusion. The method reaches state-of-the-art performances in terms of …

Face SwappingGPUImage GenerationImage Inpainting+1

SoulX-FlashTalk: Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional Distillation

2025-12-29 · Le Shen, Qian Qiao, Tan Yu, Ke Zhou 외 arxiv

Deploying massive diffusion models for real-time, infinite-duration, audio-driven avatar generation presents a significant engineering challenge, primarily due to the conflict between computational load and strict latenc…

FlashVideo:Flowing Fidelity to Detail for Efficient High-Resolution Video Generation

2025-02-07 · Shilong Zhang, Wenbo Li, Shoufa Chen, Chongjian Ge 외

DiT diffusion models have achieved great success in text-to-video generation, leveraging their scalability in model capacity and data scale. High content and motion fidelity aligned with text prompts, however, often requ…

Computational EfficiencyText-to-Video GenerationVideo Generation

FlashFace: Human Image Personalization with High-fidelity Identity Preservation

2024-03-25 · Shilong Zhang, Lianghua Huang, Xi Chen, Yifei Zhang 외

This work presents FlashFace, a practical tool with which users can easily personalize their own photos on the fly by providing one or a few reference face images and a text prompt. Our approach is distinguishable from e…

Face SwappingImage GenerationInstruction FollowingText to Image Generation+1