paper-with-me

Papers

Cross-View World Models

2026-02-07 · Rishabh Sharma, Gijs Hogervorst, Wayne E. Mackey, David J. Heeger, Stefano Martiniani arxiv

World models enable agents to plan by imagining future states, but existing approaches operate from a single viewpoint, typically egocentric, even when other perspectives would make planning easier; navigation, for instance, benefits from a bird's-eye view. We introduce Cross-View World Models (XVWM), trained with a cross-view prediction objective: given a sequence of frames from one viewpoint, predict the future state from the same or a different viewpoint after an action is taken. Enforcing cross-view consistency acts as geometric regularization: because the input and output views may share little or no visual overlap, to predict across viewpoints, the model must learn view-invariant representations of the environment's 3D structure. We train on synchronized multi-view gameplay data from Aimlabs, an aim-training platform providing precisely aligned multi-camera recordings with high-frequency action labels. The resulting model gives agents parallel imagination streams across viewpoints, enabling planning in whichever frame of reference best suits the task while executing from the egocentric view. Our results show that multi-view consistency provides a strong learning signal for spatially grounded representations. Finally, predicting the consequences of one's actions from another viewpoint may offer a foundation for perspective-taking in multi-agent settings.

📄 PDF Abstract BibTeX arXiv:2602.07277

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

2026-06-16 · Yuhang Huang, Xuan Lv, Junyan Xu, Zhiyuan Yu 외 arxiv

World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consistency required for robotic manipulation. While robotic systems rely on mult…

DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models

2023-09-18 · Zoe De Simone, Angie Boggust, Arvind Satyanarayan, Ashia Wilson

Generative text-to-image (TTI) models produce high-quality images from short textual descriptions and are widely used in academic and creative domains. Like humans, TTI models have a worldview, a conception of the world …

Fairness

WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration

2025-06-25 · Chaojun Ni, Jie Li, Haoyun Li, Hengyu Liu 외

Interactive 3D scene generation from a single image has gained significant attention due to its potential to create immersive virtual worlds. However, a key challenge in current 3D generation methods is the limited explo…

3D GenerationScene GenerationVideo Restoration

MV-TAP: Tracking Any Point in Multi-View Videos

2025-12-01 · Jahyeok Koo, Inès Hyeonsu Kim, Mungyeom Kim, Junghyun Park 외 arxiv

Multi-view camera systems enable rich observations of complex real-world scenes, and understanding dynamic objects in multi-view settings has become central to various applications. In this work, we present MV-TAP, a nov…

Point Tracking

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data

2026-06-01 · Teng Hu, Mingchun Lu, Yating Wang, Jiangning Zhang 외 arxiv

Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective. Extending these models …