paper-with-me

홈 › Papers

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization

2024-12-20 · Zihan Ding, Chi Jin, Difan Liu, Haitian Zheng, Krishna Kumar Singh, Qiang Zhang, Yan Kang, Zhe Lin, Yuchen Liu

Diffusion probabilistic models have shown significant progress in video generation; however, their computational efficiency is limited by the large number of sampling steps required. Reducing sampling steps often compromises video quality or generation diversity. In this work, we introduce a distillation method that combines variational score distillation and consistency distillation to achieve few-step video generation, maintaining both high quality and diversity. We also propose a latent reward model fine-tuning approach to further enhance video generation performance according to any specified reward metric. This approach reduces memory usage and does not require the reward to be differentiable. Our method demonstrates state-of-the-art performance in few-step generation for 10-second videos (128 frames at 12 FPS). The distilled student model achieves a score of 82.57 on VBench, surpassing the teacher model as well as baseline models Gen-3, T2V-Turbo, and Kling. One-step distillation accelerates the teacher model's diffusion sampling by up to 278.6 times, enabling near real-time generation. Human evaluations further validate the superior performance of our 4-step student models compared to teacher model using 50-step DDIM sampling.

📄 PDF Abstract BibTeX arXiv:2412.15689

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyDiversityVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

VideoLCM: Video Latent Consistency Model

2023-12-14 · Xiang Wang, Shiwei Zhang, Han Zhang, Yu Liu 외

Consistency models have demonstrated powerful capability in efficient image generation and allowed synthesis within a few sampling steps, alleviating the high computational cost in diffusion models. However, the consiste…

Computational EfficiencyImage GenerationmodelVideo Generation

Reward Lightning: Fast Video Generation via Homologous Preference Distillation

2026-07-04 · Jiaxiang Cheng, Bing Ma, Xuhua Ren, Kai Yu 외 arxiv

Achieving simultaneous preference alignment and distillation acceleration in video diffusion models remains an open challenge. Existing methods optimize the two objectives over mismatched representation spaces, where imp…

Video Generation

GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation

2026-02-02 · Xiao Liang, Yunzhu Zhang, Linchao Zhu arxiv

Diffusion models have achieved remarkable success in video generation; however, the high computational cost of the denoising process remains a major bottleneck. Existing approaches have shown promise in reducing the numb…

Computational EfficiencyVideo Generation

OSV: One Step is Enough for High-Quality Image to Video Generation

2024-09-17 · CVPR 2025 1 · Xiaofeng Mao, Zhengkai Jiang, Fu-Yun Wang, Wenbing Zhu 외

Video diffusion models have shown great potential in generating high-quality videos, making them an increasingly popular focus. However, their inherent iterative nature leads to substantial computational and time costs. …

Image to Video GenerationVideo Generation

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation

2026-04-16 · Xiangyu Liu, Feng Gao, Xiaomei Zhang, Yong Zhang 외 arxiv

Existing audio-driven video digital human generation models rely on multi-step denoising, resulting in substantial computational overhead that severely limits their deployment in real-world settings. While one-step disti…