paper-with-me

Papers

Video World Models with Long-term Spatial Memory

2025-06-05 · Tong Wu, Shuai Yang, Ryan Po, Yinghao Xu, Ziwei Liu, Dahua Lin, Gordon Wetzstein

Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal context window sizes, these models often struggle to maintain scene consistency during revisits, leading to severe forgetting of previously generated environments. Inspired by the mechanisms of human memory, we introduce a novel framework to enhancing long-term consistency of video world models through a geometry-grounded long-term spatial memory. Our framework includes mechanisms to store and retrieve information from the long-term spatial memory and we curate custom datasets to train and evaluate world models with explicitly stored 3D memory mechanisms. Our evaluations show improved quality, consistency, and context length compared to relevant baselines, paving the way towards long-term consistent world generation.

📄 PDF Abstract BibTeX arXiv:2506.05284

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Long-Context State-Space Video World Models

2025-05-26 · Ryan Po, Yotam Nitzan, Richard Zhang, Berlin Chen 외

Video diffusion models have recently shown promise for world modeling through autoregressive frame prediction conditioned on actions. However, they struggle to maintain long-term memory due to the high computational cost…

Computational EfficiencyMinecraftState Space Models

RELIC: Interactive Video World Model with Long-Horizon Memory

2025-12-03 · Yicong Hong, Yiqun Mei, Chongjian Ge, Yiran Xu 외 arxiv

A truly interactive world model requires three key ingredients: real-time long-horizon streaming, consistent spatial memory, and precise user control. However, most existing approaches address only one of these aspects i…

EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory

2025-10-01 · Jiahao Wang, Luoxin Ye, TaiMing Lu, Junfei Xiao 외 arxiv

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video genera…

3D ReconstructionVideo Generation

Towards Long-Form Spatio-Temporal Video Grounding

2026-02-26 · Xin Gu, Bing Fan, Jiali Yao, Zhipeng Zhang 외 arxiv

In real scenarios, videos can span several minutes or even hours. However, existing research on spatio-temporal video grounding (STVG), given a textual query, mainly focuses on localizing targets in short videos of tens …

Spatio-Temporal Video Grounding

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

2026-06-04 · Shiqiang Lang, Jing Liu, Haoyang He, Peiwen Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving and robotic navigation require more tha…

Autonomous DrivingSpatial Reasoning