paper-with-me

홈 › Papers

MoWorld: A Flash World Model

2026-07-07 · Team Moxin, Deyi Ji, Tianrun Chen, Xin Zhang, Jiale Yang, Qi Zhu, An Zhao, Zihao Xie, Han Wang, Xuanyi Liu, Yixiang Zhou, Pei Liu, Yi Tan, Cheng Chen, Dayi Zhu, Mingyu Wei, Hanjie Xu, Jun Liao, Siqi Li, Lingyu Lu, Hongye Fang, Hongming Tan, Youjiang Zhu, Taiyu Zhang, Zejian Li, Chaotao Ding, Lanyun Zhu, Yunhe Pan, Lingyun Sun arxiv

The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive perception, planning, and control in real-world autonomous systems. To this end, we present MoWorld, a cost-effective yet high-performance Flash World Model with an end-to-end framework spanning data generation, pre-training, distillation, and efficient inference, enabling up to 50 FPS real-time interaction with cinematic visual quality without the need of high-end GPUs. To enable large-scale real-world deployment, MoWorld jointly optimizes model capability and cost throughout the entire development pipeline. Specifically, unlike existing approaches that primarily rely on large-scale video corpora, MoWorld is built upon a scalable 3D-native data engine accumulated from our large-scale 3D vision and generative modeling pipeline, enabling the efficient construction of geometrically consistent training data across diverse real-world and synthetic environments. Based on this foundation, a curriculum cross-frame pre-training strategy for stable and scalable World Model learning, an efficient denoising-step distillation algorithm to reduce diffusion training cost, and a mixed-precision parallel inference framework for low-cost real-time deployment. MoWorld is the first real-time interactive World Model built on the Neural Processing Unit (NPU) and can achieves up to 50 FPS in such the devices, enabling practical and efficient deployment at scale. Comprehensive evaluations demonstrate that MoWorld achieves leading performance; notably, its average inference cost is only 30\%-50\% of that of existing World Models, providing a practical foundation for large-scale real-world applications of World Models. We also demonstrate diverse applications of MoWorld.

📄 PDF Abstract BibTeX arXiv:2607.06216

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

2026-08-06 · Bingyuan Wang, Baistan Zhyldyzbekov, Kunyu Feng, Zeyu Wang arxiv

Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framew…

Video Generation

Flash Lightens Gray Pixels

2019-02-27 · Yanlin Qian, Song Yan, Joni-Kristian Kämäräinen, Jiri Matas

In the real world, a scene is usually cast by multiple illuminants and herein we address the problem of spatial illumination estimation. Our solution is based on detecting gray pixels with the help of flash photography. …

Stereoscopic Flash and No-Flash Photography for Shape and Albedo Recovery

2020-06-01 · CVPR 2020 6 · Xu Cao, Michael Waechter, Boxin Shi, Ye Gao 외

We present a minimal imaging setup that harnesses both geometric and photometric approaches for shape and albedo recovery. We adopt a stereo camera and a flashlight to capture a stereo image pair and a flash/no-flash pai…

Shadow Detection

Flash-Split: 2D Reflection Removal with Flash Cues and Latent Diffusion Separation

2024-12-31 · CVPR 2025 1 · Tianfu Wang, Mingyang Xie, Haoming Cai, Sachin Shah 외

Transparent surfaces, such as glass, create complex reflections that obscure images and challenge downstream computer vision applications. We introduce Flash-Split, a robust framework for separating transmitted and refle…

Reflection Removal

Flash-Splat: 3D Reflection Removal with Flash Cues and Gaussian Splats

2024-10-03 · Mingyang Xie, Haoming Cai, Sachin Shah, Yiran Xu 외

We introduce a simple yet effective approach for separating transmitted and reflected light. Our key insight is that the powerful novel view synthesis capabilities provided by modern inverse rendering methods (e.g.,~3D G…

Inverse RenderingNovel View SynthesisReflection Removal