paper-with-me

Papers

Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving

2026-05-18 · Lijun Zhou, Hongcheng Luo, Zhenxin Zhu, Cheng Chi, Mingfei Tu, Kaixin Xiong, Lei Gong, Zhanqian Wu, Zehan Zhang, Fangzhen Li, Hao Li, Yingying Shen, Jiale He, Haohui Zhu, Shan Zhao, Kai Wang, Zhiwei Zhan, Yuechuan Pu, Kaiyuan Tan, Ruiling Yang, Xianqi Wang, Tianyi Yan, Jiawei Zhou, Lei Zhang, Jingyang Zhao, Xi Zhou, Chitian Sun, Chenming Wu, Jiong Deng, Hongwei Xie, Ming Lu, Kun Ma, Long Chen, Guang Chen, Hangjun Ye, Bing Wang, Haiyang Sun arxiv

This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward reconstruction architecture driven by sparse scene queries. WorldRec initializes structured queries in 3D space, leveraging them to aggregate cross-view, cross-temporal features, thereby naturally enforcing spatial consistency across frames and yielding compact yet high-fidelity 3D Gaussian scene representations. For world generation, we propose WorldGen, a two-stage training framework of bidirectional pretraining followed by causal fine-tuning through three progressive stages (Teacher Forcing, ODE distillation, and DMD), enabling high-quality online causal video generation in as few as 4 denoising steps. Building on both modules, we further introduce the JWM, which deeply integrates WorldRec and WorldGen to achieve synergistic gains in generation stability, cross-frame consistency, and visual fidelity, providing a solid foundation for closed-loop simulation, data synthesis, and end-to-end training in autonomous driving.

📄 PDF Abstract BibTeX arXiv:2605.18137

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingVideo Generation

Similar Papers 제목 키워드 기반

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

2026-07-13 · Xinghang Li, Jun Guo, Qiwei Li, Long Qian 외 hf

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coh…

Text-to-Image GenerationScene GenerationVideo GenerationImage Editing

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

2026-07-16 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li 외 hf

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-th…

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

2026-04-20 · Jinghui Lu, Jiayi Guan, Zhijian Huang, Jinlong Li 외 arxiv

Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes a latency cost that is prohibitive for real-time deployment. L…

Trajectory PredictionAutonomous Driving

MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving

2025-03-20 · Haiguang Wang, Daqi Liu, Hongwei Xie, Haisong Liu 외

In recent years, data-driven techniques have greatly advanced autonomous driving systems, but the need for rare and diverse training data remains a challenge, requiring significant investment in equipment and labor. Worl…

Autonomous DrivingDenoisingVideo Generation

Xiaomingbot: A Multilingual Robot News Reporter

2020-07-12 · ACL 2020 6 · Runxin Xu, Jun Cao, Mingxuan Wang, Jiaze Chen 외

This paper proposes the building of Xiaomingbot, an intelligent, multilingual and multimodal software robot equipped with four integral capabilities: news generation, news translation, news reading and avatar animation. …

ArticlesNews GenerationTranslationVoice Cloning