paper-with-me

Papers

UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control

2024-03-04 · Tian Xia, Xuweiyi Chen, Sihan Xu

Video Diffusion Models have been developed for video generation, usually integrating text and image conditioning to enhance control over the generated content. Despite the progress, ensuring consistency across frames remains a challenge, particularly when using text prompts as control conditions. To address this problem, we introduce UniCtrl, a novel, plug-and-play method that is universally applicable to improve the spatiotemporal consistency and motion diversity of videos generated by text-to-video models without additional training. UniCtrl ensures semantic consistency across different frames through cross-frame self-attention control, and meanwhile, enhances the motion quality and spatiotemporal consistency through motion injection and spatiotemporal synchronization. Our experimental results demonstrate UniCtrl's efficacy in enhancing various text-to-video models, confirming its effectiveness and universality.

📄 PDF Abstract BibTeX arXiv:2403.02332

Code (1)

XuweiyiChen/UniCtrl 공식 구현 pytorch

Tasks

DiversityVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Slicedit: Zero-Shot Video Editing With Text-to-Image Diffusion Models Using Spatio-Temporal Slices

2024-05-20 · Nathaniel Cohen, Vladimir Kulikov, Matan Kleiner, Inbar Huberman-Spiegelglas 외

Text-to-image (T2I) diffusion models achieve state-of-the-art results in image synthesis and editing. However, leveraging such pretrained models for video editing is considered a major challenge. Many existing works atte…

Image GenerationVideo Editing

DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake Detection

2024-10-31 · Fan Nie, Jiangqun Ni, Jian Zhang, Bin Zhang 외

With the advancement of deepfake generation techniques, the importance of deepfake detection in protecting multimedia content integrity has become increasingly obvious. Recently, temporal inconsistency clues have been ex…

DecoderDeepFake DetectionFace Swapping

STeP: A General and Scalable Framework for Solving Video Inverse Problems with Spatiotemporal Diffusion Priors

2025-04-10 · Bingliang Zhang, Zihui Wu, Berthy T. Feng, Yang song 외

We study how to solve general Bayesian inverse problems involving videos using diffusion model priors. While it is desirable to use a video diffusion prior to effectively capture complex temporal relationships, due to th…

DiffusionAtlas: High-Fidelity Consistent Diffusion Video Editing

2023-12-05 · Shao-Yu Chang, Hwann-Tzong Chen, Tyng-Luh Liu

We present a diffusion-based video editing framework, namely DiffusionAtlas, which can achieve both frame consistency and high fidelity in editing video object appearance. Despite the success in image editing, diffusion …

ObjectVideo Editing

Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion

2025-01-08 · Yongjia Ma, Junlin Chen, Donglin Di, Qi Xie 외

Creating high-fidelity, coherent long videos is a sought-after aspiration. While recent video diffusion models have shown promising potential, they still grapple with spatiotemporal inconsistencies and high computational…

DenoisingDiversityVideo DenoisingVideo Generation