paper-with-me

홈 › Papers

$Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation

2026-03-12 · Songlin Wei, Hongyi Jing, Boqian Li, Zhenyu Zhao, Jiageng Mao, Zhenhao Ni, Sicheng He, Jie Liu, Xiawei Liu, Kaidi Kang, Sheng Zang, Weiduo Yuan, Marco Pavone, Di Huang, Yue Wang arxiv

We introduce $Ψ_0$ (Psi-Zero), an open foundation model to address challenging humanoid loco-manipulation tasks. While existing approaches often attempt to address this fundamental problem by co-training on large and diverse human and humanoid data, we argue that this strategy is suboptimal due to the fundamental kinematic and motion disparities between humans and humanoid robots. Therefore, data efficiency and model performance remain unsatisfactory despite the considerable data volume. To address this challenge, \ours\;decouples the learning process to maximize the utility of heterogeneous data sources. Specifically, we propose a staged training paradigm with different learning objectives: First, we autoregressively pre-train a VLM backbone on large-scale egocentric human videos to acquire generalizable visual-action representations. Then, we post-train a flow-based action expert on high-quality humanoid robot data to learn precise robot joint control. Our research further identifies a critical yet often overlooked data recipe: in contrast to approaches that scale with noisy Internet clips or heterogeneous cross-embodiment robot datasets, we demonstrate that pre-training on high-quality egocentric human manipulation data followed by post-training on domain-specific real-world humanoid trajectories yields superior performance. Extensive real-world experiments demonstrate that \ours\ achieves the best performance using only about 800 hours of human video data and 30 hours of real-world robot data, outperforming baselines pre-trained on more than 10$\times$ as much data by over 40\% in overall success rate across multiple tasks. We will open-source the entire ecosystem to the community, including a data processing and training pipeline, a humanoid foundation model, and a real-time action inference engine.

📄 PDF Abstract BibTeX arXiv:2603.12263

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DemoHLM: From One Demonstration to Generalizable Humanoid Loco-Manipulation

2025-10-13 · Yuhui Fu, Feiyang Xie, Chaoyi Xu, Jing Xiong 외 arxiv

Loco-manipulation is a fundamental challenge for humanoid robots to achieve versatile interactions in human environments. Although recent studies have made significant progress in humanoid whole-body control, loco-manipu…

Embodied Chain of Action Reasoning with Multi-Modal Foundation Model for Humanoid Loco-manipulation

2025-04-13 · Yu Hao, Geeta Chandra Raju Bethala, Niraj Pudasaini, Hao Huang 외

Enabling humanoid robots to autonomously perform loco-manipulation tasks in complex, unstructured environments poses significant challenges. This entails equipping robots with the capability to plan actions over extended…

NavigateObject RearrangementSpatial Reasoning

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

2026-08-17 · Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang 외 arxiv

Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulati…

Dimensionality ReductionReinforcement Learning

SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation

2026-06-06 · Songlin Wei, Zhenhao Ni, Jie Liu, Zhenyu Zhao 외 arxiv

Humanoid foundation models are advancing faster than we can evaluate them. While real-world testing is expensive and difficult to reproduce, existing simulation benchmarks focus primarily on table-top or wheeled robots. …

Motion Planning

MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation

2026-06-08 · Jia Zheng, Teli Ma, Yudong Fan, Zifan Wang 외 arxiv

World Action Models (WAMs) couple a video dynamics prior to the policy and have shown encouraging results on tabletop manipulation, but iterative denoising over high-dimensional video-action latents leaves them too slow …