paper-with-me

홈 › Papers

ControlVideo: Training-free Controllable Text-to-Video Generation

2023-05-22 · Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, WangMeng Zuo, Qi Tian

Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temporal modeling. Besides the training burden, the generated videos also suffer from appearance inconsistency and structural flickers, especially in long video synthesis. To address these challenges, we design a \emph{training-free} framework called \textbf{ControlVideo} to enable natural and efficient text-to-video generation. ControlVideo, adapted from ControlNet, leverages coarsely structural consistency from input motion sequences, and introduces three modules to improve video generation. Firstly, to ensure appearance coherence between frames, ControlVideo adds fully cross-frame interaction in self-attention modules. Secondly, to mitigate the flicker effect, it introduces an interleaved-frame smoother that employs frame interpolation on alternated frames. Finally, to produce long videos efficiently, it utilizes a hierarchical sampler that separately synthesizes each short clip with holistic coherency. Empowered with these modules, ControlVideo outperforms the state-of-the-arts on extensive motion-prompt pairs quantitatively and qualitatively. Notably, thanks to the efficient designs, it generates both short and long videos within several minutes using one NVIDIA 2080Ti. Code is available at https://github.com/YBYBZhang/ControlVideo.

📄 PDF Abstract BibTeX arXiv:2305.13077

Code (1)

ybybzhang/controlvideo 공식 구현 pytorch

Tasks

Image GenerationText-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

ControlVideo: Conditional Control for One-shot Text-driven Video Editing and Beyond

2023-05-26 · Min Zhao, Rongzhen Wang, Fan Bao, Chongxuan Li 외

This paper presents \emph{ControlVideo} for text-driven video editing -- generating a video that aligns with a given text while preserving the structure of the source video. Building on a pre-trained text-to-image diffus…

Text-to-Video EditingVideo Editing

LooseControlVideo: Directorial Video Control using Spatial Blocking

2026-06-17 · Shariq Farooq Bhat, Niloy J. Mitra, Kalyan Sunkavalli arxiv

Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and temporal dynamics are often entangled. While existing depth-con…

Text-to-Video Generation

Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos

2023-04-03 · Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang 외

Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absence of a comprehensive dataset featuring …

Image GenerationText to Image GenerationText-to-Image GenerationText-to-Video Generation+1

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

2026-07-29 · Yuyang Huang, Yabo Chen, Wenrui Dai, Ziyang Zheng 외 arxiv

Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation acro…

Video Generation

BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models

2023-12-05 · CVPR 2024 1 · Fengyuan Shi, Jiaxi Gu, Hang Xu, Songcen Xu 외

Diffusion models have made tremendous progress in text-driven image and video generation. Now text-to-image foundation models are widely applied to various downstream image synthesis tasks, such as controllable image gen…

Image GenerationModel SelectionVideo EditingVideo Generation+1