paper-with-me

Papers

WholeBodyWAM: Generalizing Pre-trained World-Action Priors to Humanoid Loco-Manipulation via WBC-Grounded Coordination

2026-09-15 · Zhuo Li, Yiming Yao, Jim Tan, Mengjie Jing, Zhipeng Dong, Fei Chen arxiv

World Action Models (WAMs) offer a promising approach to general-purpose robot manipulation by jointly modeling visual dynamics and actions. However, most WAM studies focus on tabletop or arm-centric manipulation, while humanoid loco-manipulation remains less explored. To address this gap, we introduce WholeBodyWAM, which jointly predicts future visual dynamics, manipulation actions, and whole-body control intents for generalizable humanoid loco-manipulation. It preserves pre-trained world-action priors while grounding heterogeneous whole-body controller (WBC) semantics and coordinating whole-body behavior. Extensive experiments show that WholeBodyWAM achieves an overall simulation task success rate of 91.9%, with a 0.23 improvement in real-world out-of-distribution task progress and a 70% reduction in success-rate variance across WBCs relative to the respective baselines. These results suggest a path toward scalable humanoid whole-body intelligence by extending pre-trained world-action priors through structured WBC grounding and coordination, rather than relearning whole-body behavior from scratch. Project page: https://wholebodywam.github.io/.

📄 PDF Abstract BibTeX arXiv:2609.16644

Code (1)

BaiShuanghao/my_arXiv_daily ★ 213

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

Generative Image as Action Models

2024-07-10 · Mohit Shridhar, Yat Long Lo, Stephen James

Image-generation diffusion models have been fine-tuned to unlock new capabilities such as image-editing and novel view synthesis. Can we similarly unlock image-generation models for visuomotor control? We present GENIMA,…

Image GenerationRobot ManipulationRobot Manipulation Generalization

Exploring Real World Map Change Generalization of Prior-Informed HD Map Prediction Models

2024-06-04 · Samuel M. Bateman, Ning Xu, H. Charles Zhao, Yael Ben Shalom 외

Building and maintaining High-Definition (HD) maps represents a large barrier to autonomous vehicle deployment. This, along with advances in modern online map detection models, has sparked renewed interest in the online …

Autonomous Driving

Optimizing Diffusion Priors in Image Reconstruction from a Single Observation

2026-04-22 · Frederic Wang, Katherine L. Bouman arxiv

While diffusion priors generate high-quality posterior samples across many inverse problems, they are often trained on limited training sets or purely simulated data, thus inheriting the errors and biases of these underl…

Image ReconstructionImage Deblurring

Zero-Shot Scene Reconstruction from Single Images with Deep Prior Assembly

2024-10-21 · Junsheng Zhou, Yu-Shen Liu, Zhizhong Han

Large language and vision models have been leading a revolution in visual computing. By greatly scaling up sizes of data and model parameters, the large models learn deep priors which lead to remarkable performance in va…

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

2025-03-12 · Kechun Xu, Xunlong Xia, Kaixuan Wang, Yifei Yang 외

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features fr…

Zero-shot Generalization