paper-with-me

홈 › Papers

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

2026-06-08 · Jia Zheng, Teli Ma, Yudong Fan, Zifan Wang, Shuo Yang, Junwei Liang arxiv

World Action Models (WAMs) couple a video dynamics prior to the policy and have shown encouraging results on tabletop manipulation, but iterative denoising over high-dimensional video-action latents leaves them too slow for real-time humanoid loco-manipulation. The problem is compounded by the dominant hierarchical paradigm, in which a high-level manipulation policy controls only the upper body while a low-level controller tracks coarse base commands -- placing upper and lower body in inconsistent action spaces and reducing the legs to balance-preserving locomotion. We present MotionWAM, a real-time WAM that drives autonomous humanoid loco-manipulation from a single egocentric camera by conditioning the policy on the intermediate denoising features of a video world model. MotionWAM replaces the upper-lower split with a unified motion latent and predicts whole-body motion tokens that jointly cover locomotion, torso motion, height regulation, foot interaction, and hand manipulation in a single action space. A three-stage learning framework progressively adapts the video world model to egocentric visual dynamics and to the target humanoid embodiment. On nine real-world Unitree G1 tasks, MotionWAM runs in real time, substantially outperforms Vision-Language-Action (VLA) baselines fine-tuned on the same demonstrations by over 30% in overall success rate, and executes task-driven foot interaction that decoupled upper-lower policies cannot reach. Our results suggest that video-pretrained WAMs can be lifted from tabletop manipulation to coordinated, human-like whole-body humanoid control.

📄 PDF Abstract BibTeX arXiv:2606.09215

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bootstrapping World Models from Dynamics Models in Multimodal Foundation Models

2025-06-06 · Yifu Qiu, Yftah Ziser, Anna Korhonen, Shay B. Cohen 외

To what extent do vision-and-language foundation models possess a realistic world model (observation $\times$ action $\rightarrow$ observation) and a dynamics model (observation $\times$ observation $\rightarrow$ action)…

Weakly-supervised Learning

Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review

2025-05-26 · Matthew Lisondra, Beno Benhabib, Goldie Nejat

Rapid advancements in foundation models, including Large Language Models, Vision-Language Models, Multimodal Large Language Models, and Vision-Language-Action Models have opened new avenues for embodied AI in mobile serv…

Decision Making Under UncertaintySensor FusionVision-Language-Action

Uncovering Zero-Shot Generalization Gaps in Time-Series Foundation Models Using Real-World Videos

2025-09-30 · Lujun Li, Lama Sleem, Yiqun Wang, Yangjie Xu 외 arxiv

Recent research on time-series foundation models (TSFMs) has underscored the scarcity of real-world data, often supplemented with synthetic sources in existing datasets, whose generalizability remains however debated. As…

Zero-shot Generalization

Constrained Decoding for Safe Robot Navigation Foundation Models

2025-09-01 · Parv Kapoor, Akila Ganlath, Michael Clifford, Changliu Liu 외 arxiv

Recent advances in the development of robotic foundation models have led to promising end-to-end and general-purpose capabilities in robotic systems. Trained on vast datasets of simulated and real-world trajectories, the…

Robot Navigation

HEMM: Holistic Evaluation of Multimodal Foundation Models

2024-07-03 · Paul Pu Liang, Akshay Goindani, Talha Chafekar, Leena Mathur 외

Multimodal foundation models that can holistically process text alongside images, video, audio, and other sensory modalities are increasingly used in a variety of real-world applications. However, it is challenging to ch…