paper-with-me

홈 › Papers

Geo-Align: Video Generation Alignment via Metric Geometry Reward

2026-05-22 · Zizun Li, Haoyu Guo, Runzhe Teng, Chunhua Shen, Tong He arxiv

Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme scarcity of synchronized, multi-view real-world video data. Consequently, the prevailing paradigm often exhibits limited generalization when processing out-of-distribution real-world videos, with models struggling to accurately adhere to physical scales and camera trajectories. To bridge this gap, we propose Geo-Align, the first Reinforcement Learning framework specifically designed for camera-controlled video re-rendering. Built upon a pretrained model, we optimize the model through a scale-aware perceptual reward mechanism. Specifically, we introduce a metric 3D estimator to extract precise camera trajectories from generated videos, explicitly penalizing deviations in rotation and translation. Furthermore, we meticulously designed a data pipeline strategy based on real-world conditioning videos and target camera trajectories derived from synthetic data, eliminating the reliance on paired data. Extensive experiments demonstrate that Geo-Align consistently outperforms existing supervised learning baselines in both precise camera controllability and visual fidelity, indicating the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2605.23903

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningVideo Generation

Similar Papers 제목 키워드 기반

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

2025-07-10 · Haoyu Wu, Diankun Wu, Tianyu He, Junliang Guo 외 arxiv

Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in …

Video Generation

VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward

2026-03-27 · Zhaochong An, Orest Kupyn, Théo Uscidda, Andrea Colaco 외 arxiv

Large-scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric consistency. Prior approaches improve consistency either by augmenting the generator with additional modules or a…

Video Generation

Alignment Is All You Need For X-to-4D Generation

2026-07-02 · Qiaowei Miao, Kehan Li, Yawei Luo, Yi Yang arxiv

Generative diffusion models excel at synthesizing high-quality images, videos, and 3D content under multimodal control. However, arbitrary user-defined modality-to-4D (X-to-4D) generation remains challenging due to the h…

PLA4D: Pixel-Level Alignments for Text-to-4D Gaussian Splatting

2024-05-30 · Qiaowei Miao, JinSheng Quan, Kehan Li, Yawei Luo

Previous text-to-4D methods have leveraged multiple Score Distillation Sampling (SDS) techniques, combining motion priors from video-based diffusion models (DMs) with geometric priors from multiview DMs to implicitly gui…

3D GenerationContrastive LearningText to 3D

Initialization and Alignment for Adversarial Texture Optimization

2022-07-28 · Xiaoming Zhao, Zhizhen Zhao, Alexander G. Schwing

While recovery of geometry from image and video data has received a lot of attention in computer vision, methods to capture the texture for a given geometry are less mature. Specifically, classical methods for texture ge…

Texture Synthesis