paper-with-me

홈 › Papers

EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

2026-02-18 · Ruijie Zheng, Dantong Niu, Yuqi Xie, Jing Wang, Mengda Xu, Yunfan Jiang, Fernando Castañeda, Fengyuan Hu, You Liang Tan, Letian Fu, Trevor Darrell, Furong Huang, Yuke Zhu, Danfei Xu, Linxi Fan arxiv

Human behavior is among the most scalable sources of data for learning physical intelligence, yet how to effectively leverage it for dexterous manipulation remains unclear. While prior work demonstrates human to robot transfer in constrained settings, it is unclear whether large scale human data can support fine grained, high degree of freedom dexterous manipulation. We present EgoScale, a human to dexterous manipulation transfer framework built on large scale egocentric human data. We train a Vision Language Action (VLA) model on over 20,854 hours of action labeled egocentric human video, more than 20 times larger than prior efforts, and uncover a log linear scaling law between human data scale and validation loss. This validation loss strongly correlates with downstream real robot performance, establishing large scale human data as a predictable supervision source. Beyond scale, we introduce a simple two stage transfer recipe: large scale human pretraining followed by lightweight aligned human robot mid training. This enables strong long horizon dexterous manipulation and one shot task adaptation with minimal robot supervision. Our final policy improves average success rate by 54% over a no pretraining baseline using a 22 DoF dexterous robotic hand, and transfers effectively to robots with lower DoF hands, indicating that large scale human motion provides a reusable, embodiment agnostic motor prior.

📄 PDF Abstract BibTeX arXiv:2602.16710

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Developing Vision-Language-Action Model from Egocentric Videos

2025-09-26 · Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori arxiv

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-La…

MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos

2025-04-08 · Alexey Gavryushin, Xi Wang, Robert J. S. Malate, Chenyu Yang 외

Large-scale egocentric video datasets capture diverse human activities across a wide range of scenarios, offering rich and detailed insights into how humans interact with objects, especially those that require fine-grain…

Motus2: A Self-Evolving General World Model for Dexterous Manipulation

2026-08-31 · Hongzhe Bi, Zihao Zhou, Yihang Tang, Jingrui Pang 외 arxiv

General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output h…

Domain Adaptation

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

2026-06-20 · Yangtao Chen, Zixuan Chen, Peiyang Wang, Yong-Lu Li 외 arxiv

Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well aligned with r…

EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations

2026-06-10 · Yangcen Liu, Shuo Cheng, Xinchen Yin, Woo Chul Shin 외 arxiv

Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but directly using them for robot learning …