paper-with-me

홈 › Papers

VideoLCM: Video Latent Consistency Model

2023-12-14 · Xiang Wang, Shiwei Zhang, Han Zhang, Yu Liu, Yingya Zhang, Changxin Gao, Nong Sang

Consistency models have demonstrated powerful capability in efficient image generation and allowed synthesis within a few sampling steps, alleviating the high computational cost in diffusion models. However, the consistency model in the more challenging and resource-consuming video generation is still less explored. In this report, we present the VideoLCM framework to fill this gap, which leverages the concept of consistency models from image generation to efficiently synthesize videos with minimal steps while maintaining high quality. VideoLCM builds upon existing latent video diffusion models and incorporates consistency distillation techniques for training the latent consistency model. Experimental results reveal the effectiveness of our VideoLCM in terms of computational efficiency, fidelity and temporal consistency. Notably, VideoLCM achieves high-fidelity and smooth video synthesis with only four sampling steps, showcasing the potential for real-time synthesis. We hope that VideoLCM can serve as a simple yet effective baseline for subsequent research. The source code and models will be publicly available.

📄 PDF Abstract BibTeX arXiv:2312.09109

Code (2)

ali-vilab/VGen pytorch
ali-vilab/i2vgen-xl pytorch

Tasks

Computational EfficiencyImage GenerationmodelVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

LatentWarp: Consistent Diffusion Latents for Zero-Shot Video-to-Video Translation

2023-11-01 · Yuxiang Bao, Di Qiu, Guoliang Kang, Baochang Zhang 외

Leveraging the generative ability of image diffusion models offers great potential for zero-shot video-to-video translation. The key lies in how to maintain temporal consistency across generated video frames by image dif…

DenoisingOptical Flow EstimationTranslation

VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward

2026-03-27 · Zhaochong An, Orest Kupyn, Théo Uscidda, Andrea Colaco 외 arxiv

Large-scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric consistency. Prior approaches improve consistency either by augmenting the generator with additional modules or a…

Video Generation

Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

2025-07-17 · Yudong Jin, Sida Peng, Xuan Wang, Tao Xie 외

This paper addresses the challenge of high-fidelity view synthesis of humans with sparse-view videos as input. Previous methods solve the issue of insufficient observation by leveraging 4D diffusion models to generate vi…

DenoisingGPU

JVID: Joint Video-Image Diffusion for Visual-Quality and Temporal-Consistency in Video Generation

2024-09-21 · Hadrien Reynaud, Matthew Baugh, Mischa Dombrowski, Sarah Cechnicka 외

We introduce the Joint Video-Image Diffusion model (JVID), a novel approach to generating high-quality and temporally coherent videos. We achieve this by integrating two diffusion models: a Latent Image Diffusion Model (…

Video Generation

LatentColorization: Latent Diffusion-Based Speaker Video Colorization

2024-05-09 · Rory Ward, Dan Bigioi, Shubhajit Basak, John G. Breslin 외

While current research predominantly focuses on image-based colorization, the domain of video-based colorization remains relatively unexplored. Most existing video colorization techniques operate on a frame-by-frame basi…

Colorization