paper-with-me

홈 › Papers

Pixel-to-4D: Camera-Controlled Image-to-Video Generation with Dynamic 3D Gaussians

2026-01-02 · Melonie de Almeida, Daniela Ivanova, Tong Shi, John H. Williamson, Paul Henderson arxiv

Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence and 3D consistency in single-image-conditioned video generation. However, these methods often lack robust user controllability, such as modifying the camera path, limiting their applicability in real-world applications. Most existing camera-controlled image-to-video models struggle with accurately modeling camera motion, maintaining temporal consistency, and preserving geometric integrity. Leveraging explicit intermediate 3D representations offers a promising solution by enabling coherent video generation aligned with a given camera trajectory. Although these methods often use 3D point clouds to render scenes and introduce object motion in a later stage, this two-step process still falls short in achieving full temporal consistency, despite allowing precise control over camera movement. We propose a novel framework that constructs a 3D Gaussian scene representation and samples plausible object motion, given a single image in a single forward pass. This enables fast, camera-guided video generation without the need for iterative denoising to inject object motion into render frames. Extensive experiments on the KITTI, Waymo, RealEstate10K and DL3DV-10K datasets demonstrate that our method achieves state-of-the-art video quality and inference efficiency. The project page is available at https://melonienimasha.github.io/Pixel-to-4D-Website.

📄 PDF Abstract BibTeX arXiv:2601.00678

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationPoint Clouds

Similar Papers 제목 키워드 기반

Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories

2026-04-10 · Wonbong Jang, Shikun Liu, Soubhik Sanyal, Juan Camilo Perez 외 arxiv

Recovering camera parameters from images and rendering scenes from novel viewpoints have been treated as separate tasks in computer vision and graphics. This separation breaks down when image coverage is sparse or poses …

Video GenerationPose Estimation

World-consistent Video Diffusion with Explicit 3D Modeling

2024-12-02 · CVPR 2025 1 · Qihang Zhang, Shuangfei Zhai, Miguel Angel Bautista, Kevin Miao 외

Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with effici…

3D GenerationImage GenerationImage to 3DVideo Generation

TriMotion: Modality-Agnostic Camera Control for Video Generation

2026-06-18 · Seunghyun Shin, Jifei Song, Wooseok Jeon, Hae-Gon Jeon 외 arxiv

Camera motion control is essential for directing viewpoint changes in generative systems. However, existing methods typically condition the generation process on a single specific modality, such as explicit pose trajecto…

Video Generation

CamWorldQA: Perceptual Quality Assessment of Camera-Controlled World Video Generation

2026-08-19 · Yunhe Li, Likun Wu, Sijing Wu, Xinyu Tian 외 arxiv

Recent advances in generative video models have enabled camera-controlled world video generation, allowing models to synthesize videos under user-defined camera trajectories. However, existing video quality assessment (V…

Video Quality AssessmentVideo Generation

CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion

2025-09-24 · Chenhao Ji, Chaohui Yu, Junyao Gao, Fan Wang 외 arxiv

Recently, camera-controlled video generation has seen rapid development, offering more precise control over video generation. However, existing methods predominantly focus on camera control in perspective projection vide…

Video Generation