paper-with-me

Papers

Learning World Models for Interactive Video Generation

2025-05-28 · Taiye Chen, Xun Hu, Zihan Ding, Chi Jin

Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities due to two main challenges: compounding errors and insufficient memory mechanisms. We enhance image-to-video models with interactive capabilities through additional action conditioning and autoregressive framework, and reveal that compounding error is inherently irreducible in autoregressive video generation, while insufficient memory mechanism leads to incoherence of world models. We propose video retrieval augmented generation (VRAG) with explicit global state conditioning, which significantly reduces long-term compounding errors and increases spatiotemporal consistency of world models. In contrast, naive autoregressive generation with extended context windows and retrieval-augmented generation prove less effective for video generation, primarily due to the limited in-context learning capabilities of current video models. Our work illuminates the fundamental challenges in video world models and establishes a comprehensive benchmark for improving video generation models with internal world modeling capabilities.

📄 PDF Abstract BibTeX arXiv:2505.21996

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningRetrievalRetrieval-augmented GenerationVideo GenerationVideo Retrieval

Similar Papers 제목 키워드 기반

StableWorld: Towards Stable and Consistent Long Interactive Video Generation

2026-01-21 · Ying Yang, Zhengyao Lv, Tianlin Pan, Haofan Wang 외 arxiv

In this paper, we explore the overlooked challenge of stability and temporal consistency in interactive video generation, which synthesizes dynamic and controllable video worlds through interactive behaviors such as came…

Video Generation

GameGen-X: Interactive Open-world Game Video Generation

2024-11-01 · Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin 외

We introduce GameGen-X, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos. This model facilitates high-quality, open-domain generation by…

Text-to-Video GenerationVideo Generation

Matrix-game 2.0: An open-source real-time and streaming interactive world model

2025-08-18 · Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang 외 arxiv

Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynamics and interactive behaviors. However, existing interactive world models …

Video Generation

ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

2026-03-03 · Jiayi Zhu, Jianing Zhang, Yiying Yang, Wei Cheng 외 arxiv

This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works that lack support for unified shared world construction with multi-agent inter…

Video Generation

Sekai: A Video Dataset towards World Exploration

2025-06-18 · Zhen Li, Chuanhao Li, Xiaofeng Mao, Shaoheng Lin 외

Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration. However, existing video generation datasets are not well-suited for world exploration training a…

Video Generation