paper-with-me

홈 › Papers

GimbalDiffusion: Gravity-Aware Camera Control for Video Generation

2025-12-09 · Frédéric Fortier-Chouinard, Yannick Hold-Geoffroy, Valentin Deschaintre, Matheus Gadelha, Jean-François Lalonde arxiv

Recent progress in text-to-video generation has achieved remarkable realism, yet fine-grained control over camera motion and orientation remains elusive, especially with extreme trajectories (e.g., a 180-degree turnaround, or looking directly up or down). Existing approaches typically encode camera trajectories using relative or ambiguous representations, limiting precise geometric control and offering limited support for large rotations. We introduce GimbalDiffusion, a framework that enables camera control grounded in physical-world coordinates, using gravity as a global reference. Instead of describing motion relative to previous frames, our method defines camera trajectories in an absolute coordinate system, allowing accurate, interpretable control over camera parameters. Using panoramic 360-degree videos for training, we cover the full sphere of possible viewpoints, including combinations of extreme pitch and roll that are out-of-distribution of conventional video data. To improve camera control, we introduce null-pitch conditioning, a strategy that prevents the model from overriding camera specifications in the presence of conflicting prompt content (e.g., generating grass while the camera points toward the sky). Finally, we propose new benchmarks to evaluate gravity-aware camera-controlled video generation, assessing models' ability to generate extreme camera angles and quantify their input prompt entanglement.

📄 PDF Abstract BibTeX arXiv:2512.09112

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

2026-06-11 · Yang Zhou, Ziheng Wang, Yuqin Lu, Haofeng Liu 외 arxiv

We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the input observes only a small fraction of th…

G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing

2026-05-26 · Bharath Raj Nagoor Kani, Noah Snavely arxiv

Modern feed-forward 3D reconstruction methods like VGGT predict pixel-aligned pointmaps in camera-centric coordinate frames. However, this choice of coordinate frame is not always optimal. We propose instead to predict p…

3D Reconstruction

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

2026-06-25 · Bin Hu, Yanwen Ma, Jiehui Huang, Ziliang Zhang 외 arxiv

Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering trajectories, and physical dynamics are typic…

Gravity as a Reference for Estimating a Person's Height from Video

2019-09-05 · ICCV 2019 10 · Didier Bieler, Semih Günel, Pascal Fua, Helge Rhodin

Estimating the metric height of a person from monocular imagery without additional assumptions is ill-posed. Existing solutions either require manual calibration of ground plane and camera geometry, special cameras, or r…

World-Grounded Human Motion Recovery via Gravity-View Coordinates

2024-09-10 · Zehong Shen, Huaijin Pi, Yan Xia, Zhi Cen 외

We present a novel method for recovering world-grounded human motion from monocular video. The main challenge lies in the ambiguity of defining the world coordinate system, which varies between sequences. Previous approa…