paper-with-me

Papers

VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

2024-07-17 · Sherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace, Guocheng Qian, Michael Vasilkovsky, Hsin-Ying Lee, Chaoyang Wang, Jiaxu Zou, Andrea Tagliasacchi, David B. Lindell, Sergey Tulyakov

Modern text-to-video synthesis models demonstrate coherent, photorealistic generation of complex videos from a text description. However, most existing models lack fine-grained control over camera movement, which is critical for downstream applications related to content creation, visual effects, and 3D vision. Recently, new methods demonstrate the ability to generate videos with controllable camera poses these techniques leverage pre-trained U-Net-based diffusion models that explicitly disentangle spatial and temporal generation. Still, no existing approach enables camera control for new, transformer-based video diffusion models that process spatial and temporal information jointly. Here, we propose to tame video transformers for 3D camera control using a ControlNet-like conditioning mechanism that incorporates spatiotemporal camera embeddings based on Plucker coordinates. The approach demonstrates state-of-the-art performance for controllable video generation after fine-tuning on the RealEstate10K dataset. To the best of our knowledge, our work is the first to enable camera control for transformer-based video diffusion models.

📄 PDF Abstract BibTeX arXiv:2407.12781

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion

2026-05-25 · Ting-Hsuan Chen, Ying-Huan Chen, Tao Tu, Jie-Ying Lee 외 arxiv

Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constraints that remain challenging for perspective video generators due to th…

Scene GenerationVideo Generation

ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling

2025-12-01 · Qisen Wang, Yifan Zhao, Peisen Shen, Jialu Li 외 arxiv

Although prevailing camera-controlled video generation models can produce cinematic results, lifting them directly to the generation of 3D-consistent and high-fidelity time-synchronized multi-view videos remains challeng…

Data AugmentationVideo Generation

Taming Camera-Controlled Video Generation with Verifiable Geometry Reward

2025-12-02 · Zhaoqing Wang, Xiaobo Xia, Zhuolin Bie, Jinlin Liu 외 arxiv

Recent advances in video diffusion models have remarkably improved camera-controlled video generation, but most methods rely solely on supervised fine-tuning (SFT), leaving online reinforcement learning (RL) post-trainin…

Reinforcement LearningVideo Generation

MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis

2025-10-08 · Yihao Zhi, Chenghong Li, Hongjie Liao, Xihe Yang 외 arxiv

Recent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view synthesizers. Nevertheless, current method…

Monocular Depth EstimationNovel View SynthesisVideo GenerationPoint Clouds

ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis

2024-09-03 · Wangbo Yu, Jinbo Xing, Li Yuan, WenBo Hu 외

Despite recent advancements in neural 3D reconstruction, the dependence on dense multi-view captures restricts their broader applicability. In this work, we propose \textbf{ViewCrafter}, a novel method for synthesizing h…

3D Generation3D ReconstructionNovel View SynthesisText to 3D+1