paper-with-me

Papers

Unified Camera Positional Encoding for Controlled Video Generation

2025-12-08 · Cheng Zhang, Boying Li, Meng Wei, Yan-Pei Cao, Camilo Cruz Gambardella, Dinh Phung, Jianfei Cai arxiv

Transformers have emerged as a universal backbone across 3D perception, video generation, and world models for autonomous driving and embodied AI, where understanding camera geometry is essential for grounding visual observations in three-dimensional space. However, existing camera encoding methods often rely on simplified pinhole assumptions, restricting generalization across the diverse intrinsics and lens distortions in real-world cameras. We introduce Relative Ray Encoding, a geometry-consistent representation that unifies complete camera information, including 6-DoF poses, intrinsics, and lens distortions. To evaluate its capability under diverse controllability demands, we adopt camera-controlled text-to-video generation as a testbed task. Within this setting, we further identify pitch and roll as two components effective for Absolute Orientation Encoding, enabling full control over the initial camera orientation. Together, these designs form UCPE (Unified Camera Positional Encoding), which integrates into a pretrained video Diffusion Transformer through a lightweight spatial attention adapter, adding less than 1% trainable parameters while achieving state-of-the-art camera controllability and visual fidelity. To facilitate systematic training and evaluation, we construct a large video dataset covering a wide range of camera motions and lens types. Extensive experiments validate the effectiveness of UCPE in camera-controllable video generation and highlight its potential as a general camera representation for Transformers across future multi-view, video, and 3D tasks. Code will be available at https://github.com/chengzhag/UCPE.

📄 PDF Abstract BibTeX arXiv:2512.07237

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video GenerationAutonomous Driving

Similar Papers 제목 키워드 기반

CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation

2026-05-13 · Seonghyun Jin, Youngmin Kim, Sunwoo Park, Jong Chul Ye arxiv

Camera-conditioned video generation requires positional encoding that remains reliable under changes in camera motion, lens configuration, and scene structure. However, existing attention-level camera encodings either pr…

Video Generation

Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video

2026-05-14 · Yifan Wang, Tong He arxiv

Camera-controlled video generation has made substantial progress, enabling generated videos to follow prescribed viewpoint trajectories. However, existing methods usually learn camera-specific conditioning through camera…

Video Generation

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

2026-02-26 · Tianxing Xu, Zixuan Wang, Guangyuan Wang, Li Hu 외 arxiv

World models based on video generation demonstrate remarkable potential for simulating interactive environments yet suffer from persistent difficulties in two key areas: maintaining long-term content consistency when sce…

3D ReconstructionVideo Generation

PE-Field 4D: Video Generation Models as Canvas

2026-07-17 · Yunpeng Bai, Haoxiang Li, Qixing Huang arxiv

Diffusion Transformers have recently achieved strong performance in video generation, yet controlling scene geometry under viewpoint changes and camera motion remains challenging. In this work, we revisit the role of pos…

Video Generation

FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control

2026-02-13 · Mingzhi Sheng, Zekai Gu, Peng Li, Cheng Lin 외 arxiv

Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental disentanglement of "appearance" and "mot…

Video Generation