paper-with-me

Papers

Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training

2024-12-08 · Zhenghong Zhou, Jie An, Jiebo Luo

Precise camera pose control is crucial for video generation with diffusion models. Existing methods require fine-tuning with additional datasets containing paired videos and camera pose annotations, which are both data-intensive and computationally costly, and can disrupt the pre-trained model distribution. We introduce Latent-Reframe, which enables camera control in a pre-trained video diffusion model without fine-tuning. Unlike existing methods, Latent-Reframe operates during the sampling stage, maintaining efficiency while preserving the original model distribution. Our approach reframes the latent code of video frames to align with the input camera trajectory through time-aware point clouds. Latent code inpainting and harmonization then refine the model latent space, ensuring high-quality video generation. Experimental results demonstrate that Latent-Reframe achieves comparable or superior camera control precision and video quality to training-based methods, without the need for fine-tuning on additional datasets.

📄 PDF Abstract BibTeX arXiv:2412.06029

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint

2025-12-12 · Jiapeng Tang, Kai Li, Chengxiang Yin, Liuhao Ge 외 arxiv

We introduce FactorPortrait, a video diffusion method for controllable portrait animation that enables lifelike synthesis from disentangled control signals of facial expressions, head movement, and camera viewpoints. Giv…

Novel View Synthesis

CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback

2026-01-22 · Wenhang Ge, Guibao Shen, Jiawei Feng, Luozhou Wang 외 arxiv

Recent advances in camera-controlled video diffusion models have significantly improved video-camera alignment. However, the camera controllability still remains limited. In this work, we build upon Reward Feedback Learn…

Training-free Camera Control for Video Generation

2024-06-14 · Chen Hou, Guoqiang Wei, Yan Zeng, Zhibo Chen

We propose a training-free and robust solution to offer camera movement control for off-the-shelf video diffusion models. Unlike previous work, our method does not require any supervised finetuning on camera-annotated da…

Data AugmentationVideo Generation

CameraCtrl: Enabling Camera Control for Text-to-Video Generation

2024-04-02 · Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein 외

Controllability plays a crucial role in video generation since it allows users to create desired content. However, existing models largely overlooked the precise control of camera pose that serves as a cinematic language…

Text-to-Video GenerationVideo Generation

ObjCtrl-2.5D: Training-free Object Control with Camera Poses

2024-12-10 · Zhouxia Wang, Yushi Lan, Shangchen Zhou, Chen Change Loy

This study aims to achieve more precise and versatile object control in image-to-video (I2V) generation. Current methods typically represent the spatial movement of target objects with 2D trajectories, which often fail t…

Object