paper-with-me

홈 › Papers

SF-V: Single Forward Video Generation Model

2024-06-06 · Zhixing Zhang, Yanyu Li, Yushu Wu, Yanwu Xu, Anil Kag, Ivan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Junli Cao, Dimitris Metaxas, Sergey Tulyakov, Jian Ren

Diffusion-based video generation models have demonstrated remarkable success in obtaining high-fidelity videos through the iterative denoising process. However, these models require multiple denoising steps during sampling, resulting in high computational costs. In this work, we propose a novel approach to obtain single-step video generation models by leveraging adversarial training to fine-tune pre-trained video diffusion models. We show that, through the adversarial training, the multi-steps video diffusion model, i.e., Stable Video Diffusion (SVD), can be trained to perform single forward pass to synthesize high-quality videos, capturing both temporal and spatial dependencies in the video data. Extensive experiments demonstrate that our method achieves competitive generation quality of synthesized videos with significantly reduced computational overhead for the denoising process (i.e., around $23\times$ speedup compared with SVD and $6\times$ speedup compared with existing works, with even better generation quality), paving the way for real-time video synthesis and editing. More visualization results are made publicly available at https://snap-research.github.io/SF-V.

📄 PDF Abstract BibTeX arXiv:2406.04324

Code (1)

snap-research/SF-V 공식 구현

Tasks

DenoisingmodelVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation

2026-06-17 · Lin Zhang, Sicheng Mo, Zefan Cai, Jinhong Lin 외 arxiv

Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal gener…

Story GenerationVideo Generation

SinFusion: Training Diffusion Models on a Single Image or Video

2022-11-21 · Yaniv Nikankin, Niv Haim, Michal Irani

Diffusion models exhibited tremendous progress in image and video generation, exceeding GANs in quality and diversity. However, they are usually trained on very large datasets and are not naturally adapted to manipulate …

DiversityImage ManipulationVideo Generation

Fast Multi-view Consistent 3D Editing with Video Priors

2025-11-28 · Liyi Chen, Ruihuang Li, Guowen Zhang, Pengfei Wang 외 arxiv

Text-driven 3D editing enables user-friendly 3D object or scene editing with text instructions. Due to the lack of multi-view consistency priors, existing methods typically resort to employing 2D generation or editing mo…

Video Generation

Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models

2025-11-01 · Panwang Pan, Chenguo Lin, Jingjing Zhao, Chenxin Li 외 arxiv

We introduce Diff4Splat, a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion const…

Dynamic ReconstructionNovel View SynthesisScene GenerationVideo Generation

EgoSampling: Wide View Hyperlapse from Egocentric Videos

2016-04-26 · Tavi Halperin, Yair Poleg, Chetan Arora, Shmuel Peleg

The possibility of sharing one's point of view makes use of wearable cameras compelling. These videos are often long, boring and coupled with extreme shake, as the camera is worn on a moving person. Fast forwarding (i.e.…