paper-with-me

Papers

AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video

2026-09-13 · Jiaming Tan, Mingliang Zhai, Zhen Li, Yuwei Wu, Chuanhao Li, Kaipeng Zhang hf

Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatial coverage is typically obtained by synthesizing full-sphere videos or constructing explicit 3D representations. Motivated by the complementary roles of global context and selective local acuity in visual perception, we present AlayaVista, a camera-controllable streaming video world model that decouples panoramic world evolution from perspective observation synthesis. Given a single perspective image, AlayaVista constructs a 360-degree scene prior using a pretrained panorama expansion model and then evolves the scene as a camera-conditioned panoramic latent state. A latent viewport renderer maps this state to the requested perspective video latents, while a perspective refiner restores details, suppresses artifacts, and performs super-resolution. To support efficient streaming, we adapt the panoramic generator to chunk-autoregressive generation and distill both panoramic generation and perspective refinement into few-step processes. To provide the supervision required by this design, we construct MUGEN, a large-scale real-world panoramic video dataset containing 1,318 hours of videos at resolutions of at least 4K, together with rich semantic and geometric annotations.

📄 PDF Abstract BibTeX arXiv:2609.14462

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PanoWorld: Geometry-Consistent Panoramic Video World Modeling

2026-05-14 · Le Jiang, Xiangyu Bai, Bishoy Galoaa, Shayda Moezzi 외 arxiv

We present PanoWorld, a panoramic video world model that generates geometry-consistent 360$\degree$ video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do n…

Video Generation

MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold

2026-06-11 · Yang Zhou, Ziheng Wang, Yuqin Lu, Haofeng Liu 외 arxiv

We present MoVerse, a real-time video world model that creates an interactively navigable scene from a single narrow-field-of-view image. This setting is challenging because the input observes only a small fraction of th…

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

2026-07-23 · Sicheng Mo, Yuheng Li, Ziyang Leng, Krishna Kumar Singh 외 hf

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines …

Video Generation

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

2025-07-07 · Meng Wei, Chenyang Wan, Xiqian Yu, Tai Wang 외 arxiv

Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low latency grounded in language instructions. While Video-based Large Language M…

Computational Efficiency

Big Data Meet Cyber-Physical Systems: A Panoramic Survey

2018-10-29 · Rachad Atat, Lingjia Liu, Jinsong Wu, Guangyu Li 외

The world is witnessing an unprecedented growth of cyber-physical systems (CPS), which are foreseen to revolutionize our world {via} creating new services and applications in a variety of sectors such as environmental mo…

Survey