paper-with-me

Papers

Motus: A Unified Latent Action World Model

2025-12-15 · Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, Hongyan Zhao, Hanyu Liu, Zhizhong Su, Lei Ma, Hang Su, Jun Zhu arxiv

While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinders learning from large-scale, heterogeneous data. In this paper, we propose Motus, a unified latent action world model that leverages existing general pretrained models and rich, sharable motion information. Motus introduces a Mixture-of-Transformer (MoT) architecture to integrate three experts (i.e., understanding, video generation, and action) and adopts a UniDiffuser-style scheduler to enable flexible switching between different modeling modes (i.e., world models, vision-language-action models, inverse dynamics models, video generation models, and video-action joint prediction models). Motus further leverages the optical flow to learn latent actions and adopts a recipe with three-phase training pipeline and six-layer data pyramid, thereby extracting pixel-level "delta action" and enabling large-scale action pretraining. Experiments show that Motus achieves superior performance against state-of-the-art methods in both simulation (a +15% improvement over X-VLA and a +45% improvement over Pi0.5) and real-world scenarios(improved by +11~48%), demonstrating unified modeling of all functionalities and priors significantly benefits downstream robotic tasks.

📄 PDF Abstract BibTeX arXiv:2512.13030

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Motus2: A Self-Evolving General World Model for Dexterous Manipulation

2026-08-31 · Hongzhe Bi, Zihao Zhou, Yihang Tang, Jingrui Pang 외 arxiv

General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output h…

Domain Adaptation

Motubrain: An Advanced World Action Model for Robot Control

2026-04-30 · MotuBrain Team, Chendong Xiang, Fan Bao, Haitian Liu 외 arxiv

Vision-Language-Action (VLA) models generalize semantically well but often lack fine-grained modeling of world dynamics. We present Motubrain, a unified World Action Model that jointly models video and action under a Uni…

Video Generation

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

2026-08-11 · Xiao Liu, Yuguang Yang, Xi Wang, Kai Jiang 외 arxiv

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk…

Robot Manipulation

Numerical Approximation Methods for Antenna Radiation Patterns for Motus Wildlife Tracking Systems

2022-06-28 · Erik Carlson, Douglas Gobeille, Robert Deluca, Pam Loring

As plans for offshore wind energy development increase in the US, the developing methods to monitor migratory birds and bats offshore is an important area of research. To contribute to this research, current guidance rec…

IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI

2024-10-17 · Xiaoyu Chen, Junliang Guo, Tianyu He, Chuheng Zhang 외

We introduce Image-GOal Representations (IGOR), aiming to learn a unified, semantically consistent action space across human and various robots. Through this unified latent action space, IGOR enables knowledge transfer a…

Transfer Learning