paper-with-me

Papers

PanoWorld: Geometry-Consistent Panoramic Video World Modeling

2026-05-14 · Le Jiang, Xiangyu Bai, Bishoy Galoaa, Shayda Moezzi, Caleb James Lee, Tooba Imtiaz, Edmund Yeh, Jennifer Dy, Yanzhi Wang, Sarah Ostadabbas arxiv

We present PanoWorld, a panoramic video world model that generates geometry-consistent 360$\degree$ video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do not explicitly constrain the underlying 3D scene state, producing outputs that appear plausible yet exhibit inconsistent depth, broken correspondences, and implausible motion across the spherical surface. We address this gap by framing panoramic video generation as a geometry- and dynamics-consistent latent state modeling problem rather than pure visual synthesis. Building on a pre-trained perspective video world model, we introduce two lightweight regularizers: a depth consistency loss against pseudo ground-truth panoramic depth, and a trajectory consistency loss that supervises the 3D world-frame positions of tracked points across time. We further apply spherical-geometry-aware adaptation to the conditioning and positional encoding. We additionally introduce PanoGeo, a unified geometry-aware panoramic video dataset with consistent depth, trajectory, and prompt annotations across diverse real and synthetic sources, used for both training and stratified evaluation. Experiments show that PanoWorld improves geometric consistency over prior panoramic generation methods while maintaining competitive visual realism, establishing that panoramic video generation must be treated as a geometric modeling problem to support the holistic spatial understanding requirements of embodied AI applications. Code is available at https://github.com/ostadabbas/PanoWorld.

📄 PDF Abstract BibTeX arXiv:2605.15391

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation

2026-07-29 · Yongxin Su, Linjie Hou, Feng Wang, Jialin Tang 외 arxiv

We address the problem of reconstructing a high-fidelity, freely navigable 3D scene from a single $360^\circ$ panorama, without per-scene optimization or multi-view capture. Existing methods either lack metric trajectory…

3D ReconstructionVideo Generation

PanoWorld: Real-World Panoramic Generation

2026-07-10 · Haoyuan Li, Dizhe Zhang, Yuemei Zhou, Xiangkai Zhang 외 arxiv

In this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotation-equivariant property of omnidirectional representations, where rotation can be treated as an implici…

PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion

2025-09-29 · Yuyang Yin, HaoXiang Guo, Fangfu Liu, Mengyu Wang 외 arxiv

Generating a complete and explorable 360-degree visual world enables a wide range of downstream applications. While prior works have advanced the field, they remain constrained by either narrow field-of-view limitations,…

Video Generation

PanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama Synthesis

2026-05-18 · Jinrang Jia, Zhenjia Li, Yijiang Hu, Yifeng Shi arxiv

Generating a consistent whole-house VR tour from a floorplan and style reference requires both photorealistic panoramas and cross-view spatial coherence. Pure 2D generators produce appealing single panoramas but re-imagi…

3D Generation

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World

2026-05-13 · Changpeng Wang, Xin Lin, Junhan Liu, Yuheng Liu 외 arxiv

Multimodal large laboratory models (MLLMs) still struggle with spatial understanding under the dominant perspective-image paradigm, which inherits the narrow field of view of human-like perception. For navigation, roboti…

Scene UnderstandingSpatial Reasoning