paper-with-me

홈 › Papers

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

2026-09-21 · Wangbo Yu, Kunhao Liu, Wenbo Hu, Shenghai Yuan, Chaoran Feng, Haiyang Zhou, Yukun Huang, Yiran Wang, Wang Zhao, Yingmin Luo, Ying Shan hf

Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.

📄 PDF Abstract BibTeX arXiv:2609.24984

Code (5)

BaiShuanghao/my_arXiv_daily ★ 213
InsomaniacElf/sg-tamil-tts-resources- ★ 1
TencentARC/WorldCrafter ★ 238
🤗 TencentARC/WorldCrafter-Base ★ 3
🤗 TencentARC/WorldCrafter-Fast ★ 7

Similar Papers 제목 키워드 기반

FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction

2025-09-25 · Yixiang Dai, Fan Jiang, Chiyu Wang, Mu Xu 외 arxiv

High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established stron…

Novel View SynthesisVideo Generation

Geometry-Aware Implicit Memory for Video World Models

2026-06-01 · Zhengxuan Wei, Xu Guo, Xinghui Li, Xunzhi Xiang 외 arxiv

Video world models aim to simulate controllable visual environments, but long-horizon rollouts depend on what the model remembers after observations leave its native context window. Explicit memories retain frames or onl…

I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation

2026-03-24 · Jia Li, Han Yan, Yihang Chen, Siqi Li 외 arxiv

Despite remarkable progress in video generation, maintaining long-term scene consistency upon revisiting previously explored areas remains challenging. Existing solutions rely either on explicitly constructing 3D geometr…

Novel View Synthesis3D ReconstructionScene GenerationVideo Generation

CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation

2026-02-06 · Kaiyi Huang, Yukun Huang, Yu Li, Jianhong Bai 외 arxiv

Cinematic video production requires control over scene-subject composition and camera movement, but live-action shooting remains costly due to the need for constructing physical sets. To address this, we introduce the ta…

Text-to-Video Generation

DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-based Human Video Generation

2024-01-01 · CVPR 2024 1 · Chenyang Wang, Zerong Zheng, Tao Yu, Xiaoqian Lv 외

Existing diffusion models for pose-guided human video generation mostly suffer from temporal inconsistency in the generated appearance and poses due to the inherent randomization nature of the generation process. In …

Video Generation