paper-with-me

홈 › Papers

ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation

2026-05-07 · Omar El Khalifi, Thomas Rossi, Oscar Fossey, Thibault Fouque, Ulysse Mizrahi, Philip Torr, Ivan Laptev, Fabio Pizzati, Baptiste Bellot-Gurlet arxiv

For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two-phase conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose-only guidance refines high-frequency details without over-constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose-only control and other pose and camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations, especially under large viewpoint changes. Our results highlight that careful camera-consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/.

📄 PDF Abstract BibTeX arXiv:2605.06667

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Platypose: Calibrated Zero-Shot Multi-Hypothesis 3D Human Motion Estimation

2024-03-10 · Paweł A. Pierzchlewicz, Caio O. da Silva, R. James Cotton, Fabian H. Sinz

Single camera 3D pose estimation is an ill-defined problem due to inherent ambiguities from depth, occlusion or keypoint noise. Multi-hypothesis pose estimation accounts for this uncertainty by providing multiple 3D pose…

3D Pose EstimationMotion EstimationPose Estimation

MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

2026-09-04 · Zijie Zhu, Weiren Cai, Yizhou Wang, Zhenjie Yang 외 arxiv

Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity understanding, robot learning, and augmented reality. Existing systems typically decompose this problem into s…

CamMimic: Zero-Shot Image To Camera Motion Personalized Video Generation Using Diffusion Models

2025-04-13 · Pooja Guhan, Divya Kothandaraman, Tsung-Wei Huang, Guan-Ming Su 외

We introduce CamMimic, an innovative algorithm tailored for dynamic video editing needs. It is designed to seamlessly transfer the camera motion observed in a given reference video onto any scene of the user's choice in …

Video EditingVideo Generation

Zero-Shot Metric Depth with a Field-of-View Conditioned Diffusion Model

2023-12-20 · Saurabh Saxena, Junhwa Hur, Charles Herrmann, Deqing Sun 외

While methods for monocular depth estimation have made significant strides on standard benchmarks, zero-shot metric depth estimation remains unsolved. Challenges include the joint modeling of indoor and outdoor scenes, w…

DenoisingDepth EstimationMonocular Depth Estimation

WristCompass: Kinematic Coupling as a Learnable Visual Concept for Ego-Camera Orientation

2026-05-29 · Varun Nair, Vidyut Baradwaj, Jiahang He, Anya Singh 외 arxiv

Recovering ego-camera orientation from manipulation video is a prerequisite for disentangling hand motion from camera motion, a key step in imitation learning from egocentric demonstrations. The obvious approach, inferri…