paper-with-me

Papers

OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation

2026-03-31 · Yuheng Liu, Xin Lin, Xinke Li, Baihan Yang, Chen Wang, Kalyan Sunkavalli, Yannick Hold-Geoffroy, Hao Tan, Kai Zhang, Xiaohui Xie, Zifan Shi, Yiwei Hu arxiv

Modeling scenes using video generation models has garnered growing research interest in recent years. However, most existing approaches rely on perspective video models that synthesize only limited observations of a scene, leading to issues of completeness and global consistency. We propose OmniRoam, a controllable panoramic video generation framework that exploits the rich per-frame scene coverage and inherent long-term spatial and temporal consistency of panoramic representation, enabling long-horizon scene wandering. Our framework begins with a preview stage, where a trajectory-controlled video generation model creates a quick overview of the scene from a given input image or video. Then, in the refine stage, this video is temporally extended and spatially upsampled to produce long-range, high-resolution videos, thus enabling high-fidelity world wandering. To train our model, we introduce two panoramic video datasets that incorporate both synthetic and real-world captured videos. Experiments show that our framework consistently outperforms state-of-the-art methods in terms of visual quality, controllability, and long-term scene consistency, both qualitatively and quantitatively. We further showcase several extensions of this framework, including real-time video generation and 3D reconstruction. Code is available at https://github.com/yuhengliu02/OmniRoam.

📄 PDF Abstract BibTeX arXiv:2603.30045

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionVideo Generation

Similar Papers 제목 키워드 기반

Think before Go: Hierarchical Reasoning for Image-goal Navigation

2026-04-19 · Pengna Li, Kangyi Wu, Shaoqing Xu, Fang Li 외 arxiv

Image-goal navigation steers an agent to a target location specified by an image in unseen environments. Existing methods primarily handle this task by learning an end-to-end navigation policy, which compares the similar…

Reinforcement Learning

EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory

2025-10-01 · Jiahao Wang, Luoxin Ye, TaiMing Lu, Junfei Xiao 외 arxiv

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video genera…

3D ReconstructionVideo Generation

OmniNWM: Omniscient Driving Navigation World Models

2025-10-21 · Bohan Li, Zhuang Ma, Dalong Du, Baorui Peng 외 arxiv

Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. However, existing methods are typically restricted to fragmented modality modeling, short-horizon …

Autonomous Driving

CylindTrack: Depth-Aware Cylindrical Motion Modeling for Panoramic Multi-Object Tracking

2026-06-29 · Buyin Deng, Kai Luo, Lingxin Huang, Xinqi Liu 외 arxiv

Multi-Object Tracking (MOT) is a core capability for embodied perception, and panoramic cameras are attractive for embodied systems because their 360° field of view reduces blind spots and keeps surrounding targets obser…

Multi-Object TrackingTrajectory Modeling

AtlantaNet: Inferring the 3D Indoor Layout from a Single 360(∘) Image beyond the Manhattan World Assumption

2020-08-01 · ECCV 2020 8 · Giovanni Pintore, Marco Agus, Enrico Gobbetti

We introduce a novel end-to-end approach to predict a 3D room layout from a single panoramic image. Compared to recent state-of-the-art works, our method is not limited to Manhattan World environments, and can reconstruc…

3D Room Layouts From A Single RGB PanoramaDecoder