paper-with-me

Papers

Yume-1.5: A Text-Controlled Interactive World Generation Model

2025-12-26 · Xiaofeng Mao, Zhen Li, Chuanhao Li, Xiaojie Xu, Kaining Ying, Tong He, Jiangmiao Pang, Yu Qiao, Kaipeng Zhang arxiv

Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy inference steps, and rapidly growing historical context, which severely limit real-time performance and lack text-controlled generation capabilities. To address these challenges, we propose \method, a novel framework designed to generate realistic, interactive, and continuous worlds from a single image or text prompt. \method achieves this through a carefully designed framework that supports keyboard-based exploration of the generated worlds. The framework comprises three core components: (1) a long-video generation framework integrating unified context compression with linear attention; (2) a real-time streaming acceleration strategy powered by bidirectional attention distillation and an enhanced text embedding scheme; (3) a text-controlled method for generating world events. We have provided the codebase in the supplementary material.

📄 PDF Abstract BibTeX arXiv:2512.22096

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Yume: An Interactive World Generation Model

2025-07-23 · Xiaofeng Mao, Shaoheng Lin, Zhen Li, Chuanhao Li 외 arxiv

Yume aims to use images, text, or videos to create an interactive, realistic, and dynamic world, which allows exploration and control using peripheral devices or neural signals. In this report, we present a preview versi…

Video Generation

Sekai: A Video Dataset towards World Exploration

2025-06-18 · Zhen Li, Chuanhao Li, Xiaofeng Mao, Shaoheng Lin 외

Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration. However, existing video generation datasets are not well-suited for world exploration training a…

Video Generation

WorldMark: A Unified Benchmark Suite for Interactive Video World Models

2026-04-23 · Xiaojie Xu, Zhengyuan Lin, Kang He, Yukang Feng 외 arxiv

Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated on its own benchmark with private scenes and trajectories, making fair cross-model co…

Video Generation

BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression

2026-06-08 · Shaohao Rui, Xiaofeng Mao, Zhanyu Zhang, Peijia Lin 외 arxiv

Transitioning bidirectional video diffusion models into an autoregressive paradigm improves the interactivity of video world models, but existing causal pipelines need many stages (control fine-tuning, autoregressive tra…

LaGen: Towards Autoregressive LiDAR Scene Generation

2025-11-26 · Sizhuo Zhou, Xiaosong Jia, Fanrui Zhang, Junjie Li 외 arxiv

Generative world models for autonomous driving (AD) are of great value in applications such as data augmentation, closed-loop simulation, and safety-critical scenario evaluation. Unlike the widely studied image modality,…

Autonomous DrivingData AugmentationScene Generation