paper-with-me

Papers

Efficient Video Diffusion Models: Advancements and Challenges

2026-04-17 · Shitong Shao, Lichen Bai, Pengfei Wan, James Kwok, Zeke Xie arxiv

Video diffusion models have rapidly become the dominant paradigm for high-fidelity generative video synthesis, but their practical deployment remains constrained by severe inference costs. Compared with image generation, video synthesis compounds computation across spatial-temporal token growth and iterative denoising, making attention and memory traffic major bottlenecks in real-world settings. This survey provides a systematic and deployment-oriented review of efficient video diffusion models. We propose a unified categorization that organizes existing methods into four classes of main paradigms, including step distillation, efficient attention, model compression, and cache/trajectory optimization. Building on this categorization, we respectively analyze algorithmic trends of these four paradigms and examine how different design choices target two core objectives: reducing the number of function evaluations and minimizing per-step overhead. Finally, we discuss open challenges and future directions, including quality preservation under composite acceleration, hardware-software co-design, robust real-time long-horizon generation, and open infrastructure for standardized evaluation. To the best of our knowledge, our work is the first comprehensive survey on efficient video diffusion models, offering researchers and engineers a structured overview of the field and its emerging research directions.

📄 PDF Abstract BibTeX arXiv:2604.15911

Code (0)

등록된 구현이 없습니다.

Tasks

Model CompressionImage Generation

Similar Papers 제목 키워드 기반

Multi-scale 2D Temporal Map Diffusion Models for Natural Language Video Localization

2024-01-16 · Chongzhi Zhang, Mingyuan Zhang, Zhiyang Teng, Jiayi Li 외

Natural Language Video Localization (NLVL), grounding phrases from natural language descriptions to corresponding video segments, is a complex yet critical task in video understanding. Despite ongoing advancements, many …

DecoderDenoisingVideo Understanding

RelightVid: Temporal-Consistent Diffusion Model for Video Relighting

2025-01-27 · Ye Fang, Zeyi Sun, Shangzhan Zhang, Tong Wu 외

Diffusion models have demonstrated remarkable success in image generation and editing, with recent advancements enabling albedo-preserving image relighting. However, applying these models to video relighting remains chal…

Image GenerationImage Relightingmodel

OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization

2024-12-19 · Jiacheng Zhang, Jie Wu, Weifeng Chen, Yatai Ji 외

In recent years, the field of text-to-video (T2V) generation has made significant strides. Despite this progress, there is still a gap between theoretical advancements and practical application, amplified by issues like …

Video Quality AssessmentVisual Question Answering (VQA)

ANYPORTAL: Zero-Shot Consistent Video Background Replacement

2025-09-09 · Wenshuo Gao, Xicheng Lan, Shuai Yang arxiv

Despite the rapid advancements in video generation technology, creating high-quality videos that precisely align with user intentions remains a significant challenge. Existing methods often fail to achieve fine-grained c…

Video Generation

ReVision: High-Quality, Low-Cost Video Generation with Explicit 3D Physics Modeling for Complex Motion and Interaction

2025-04-30 · Qihao Liu, Ju He, Qihang Yu, Liang-Chieh Chen 외

In recent years, video generation has seen significant advancements. However, challenges still persist in generating complex motions and interactions. To address these challenges, we introduce ReVision, a plug-and-play f…

Video Generation