paper-with-me

홈 › Papers

Spatia: Video Generation with Updatable Spatial Memory

2025-12-17 · Jinjing Zhao, Fangyun Wei, Zhening Liu, Hongyang Zhang, Chang Xu, Yan Lu arxiv

Existing video generation models struggle to maintain long-term spatial and temporal consistency due to the dense, high-dimensional nature of video signals. To overcome this limitation, we propose Spatia, a spatial memory-aware video generation framework that explicitly preserves a 3D scene point cloud as persistent spatial memory. Spatia iteratively generates video clips conditioned on this spatial memory and continuously updates it through visual SLAM. This dynamic-static disentanglement design enhances spatial consistency throughout the generation process while preserving the model's ability to produce realistic dynamic entities. Furthermore, Spatia enables applications such as explicit camera control and 3D-aware interactive editing, providing a geometrically grounded framework for scalable, memory-driven video generation.

📄 PDF Abstract BibTeX arXiv:2512.15716

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation

2026-04-20 · Yanjun Guo, Zhengqiang Zhang, Pengfei Wang, Xinyue Liang 외 arxiv

Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory modeling with video generation, leading t…

Video Generation

EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory

2025-10-01 · Jiahao Wang, Luoxin Ye, TaiMing Lu, Junfei Xiao 외 arxiv

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video genera…

3D ReconstructionVideo Generation

EgoSim: Egocentric World Simulator for Embodied Interaction Generation

2026-04-01 · Jinkun Hao, Mingda Jia, Ruiyan Wang, Hongrui Zhu 외 arxiv

We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation. Existing egocentric s…

Point Clouds

Latent Spatial Memory for Video World Models

2026-06-08 · Weijie Wang, Haoyu Zhao, Yifan Yang, Feng Chen 외 arxiv

Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computationally expensive, requiring repeated re…

Video Generation

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

2026-06-04 · Shiqiang Lang, Jing Liu, Haoyang He, Peiwen Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving and robotic navigation require more tha…

Autonomous DrivingSpatial Reasoning