paper-with-me

Papers

CameraCtrl: Enabling Camera Control for Text-to-Video Generation

2024-04-02 · Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, Ceyuan Yang

Controllability plays a crucial role in video generation since it allows users to create desired content. However, existing models largely overlooked the precise control of camera pose that serves as a cinematic language to express deeper narrative nuances. To alleviate this issue, we introduce CameraCtrl, enabling accurate camera pose control for text-to-video(T2V) models. After precisely parameterizing the camera trajectory, a plug-and-play camera module is then trained on a T2V model, leaving others untouched. Additionally, a comprehensive study on the effect of various datasets is also conducted, suggesting that videos with diverse camera distribution and similar appearances indeed enhance controllability and generalization. Experimental results demonstrate the effectiveness of CameraCtrl in achieving precise and domain-adaptive camera control, marking a step forward in the pursuit of dynamic and customized video storytelling from textual and camera pose inputs. Our project website is at: https://hehao13.github.io/projects-CameraCtrl/.

📄 PDF Abstract BibTeX arXiv:2404.02101

Code (1)

hehao13/cameractrl 공식 구현 pytorch

Tasks

Text-to-Video GenerationVideo Generation

Similar Papers 제목 키워드 기반

CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models

2025-03-13 · Hao He, Ceyuan Yang, Shanchuan Lin, Yinghao Xu 외

This paper introduces CameraCtrl II, a framework that enables large-scale dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from dimin…

CamCtrl3D: Single-Image Scene Exploration with Precise 3D Camera Control

2025-01-10 · Stefan Popov, Amit Raj, Michael Krainin, Yuanzhen Li 외

We propose a method for generating fly-through videos of a scene, from a single image and a given camera trajectory. We build upon an image-to-video latent diffusion model. We condition its UNet denoiser on the camera tr…

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

2026-03-16 · Minjun Kang, Inkyu Shin, Taeyeop Lee, Myungchul Kim 외 arxiv

Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, …

Novel View Synthesis

TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

2026-07-30 · Jiwen Liu, Shujuan Li, Xiaohan Li, Zijie Meng 외 arxiv

Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizi…

3D Reconstruction

OmniCam: Unified Multimodal Video Generation via Camera Control

2025-04-03 · Xiaoda Yang, Jiayang Xu, Kaixuan Luan, Xinyu Zhan 외

Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as complex interaction and limited control ca…

Video Generation