paper-with-me

홈 › Papers

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

2026-05-18 · Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo arxiv

Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly to obtain and tightly coupled to specific embodiments. Human videos, by contrast, are abundant and capture rich interactions, providing diverse semantic and physical cues for real-world manipulation. Yet, embodiment differences and the frequent absence of task-aligned annotations make their direct use in VLA models challenging. This survey provides a unified view of how human videos are transformed into effective knowledge for VLA models. We categorize existing approaches into four classes based on the action-related information they derive: (i) latent action representations that encode inter-frame changes; (ii) predictive world models that forecast future frames; (iii) explicit 2D supervision that extracts image-plane cues; and (iv) explicit 3D reconstruction that recovers geometry or motion. Beyond this taxonomy, we highlight three key open challenges in this area: structuring unstructured videos into training-ready episodes, grounding video-derived supervision into robot-executable actions under embodiment and viewpoint heterogeneity, and designing evaluation protocols that better predict real-world deployment performance and transfer efficiency, thereby informing future research directions. A curated list of papers and resources is available at https://github.com/AaronFengZY/HumanCentricToVLA-Survey.

📄 PDF Abstract BibTeX arXiv:2606.00054

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation3D Reconstruction

Similar Papers 제목 키워드 기반

Learning by Watching: A Review of Video-based Learning Approaches for Robot Manipulation

2024-02-11 · Chrisantus Eze, Christopher Crick

Robot learning of manipulation skills is hindered by the scarcity of diverse, unbiased datasets. While curated datasets can help, challenges remain in generalizability and real-world transfer. Meanwhile, large-scale "in-…

Representation LearningRobot ManipulationSurvey

EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations

2026-06-10 · Yangcen Liu, Shuo Cheng, Xinchen Yin, Woo Chul Shin 외 arxiv

Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but directly using them for robot learning …

Robot Learning from Human Videos: A Survey

2026-04-30 · Junyi Ma, Erhang Zhang, Haoran Yang, Ditao Li 외 arxiv

A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of learning robot manipulation skills from human video data has attracted…

Robot ManipulationVideo Generation

RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience

2026-08-19 · Yaowei Guo, Zeng Tao, Yuxin Jiang, Yunuo Chen 외 arxiv

Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transfo…

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

2026-06-20 · Yangtao Chen, Zixuan Chen, Peiyang Wang, Yong-Lu Li 외 arxiv

Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well aligned with r…