paper-with-me

홈 › Papers

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

2026-09-18 · Jinho Jeong, Se June Joo, Jaehyun Kang, Dongyun Kim, Yena Kim, Hanjung Kim, Seon Joo Kim hf

Human video datasets offer an abundant and diverse source of interaction data that can complement expensive real-robot data. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately at scale. In this work, we systematically examine whether robotized human videos can serve as an effective and scalable source of supervision for VLA pretraining. To this end, we develop a robotization pipeline that converts heterogeneous human videos into robot-aligned observations and action trajectories while inferring missing intermediate signals across annotation levels. Using this pipeline, we construct the HuRo dataset, comprising about 630K robotized episodes and 142M processed frames from five human-video sources. Across four real-world manipulation tasks, increasing the amount of robotized pretraining data improves overall completion from 51.5% to 80.3% and OOD completion under spatial and visual shifts from 34.9% to 72.2%. Ablations further show that visual robotization improves OOD robustness and that end-to-end pretraining with retargeted actions outperforms visual-only transfer. Project website: https://3587jjh.github.io/HuRo.

📄 PDF Abstract BibTeX arXiv:2609.10706

Code (1)

3587jjh/HuRo ★ 43

Similar Papers 제목 키워드 기반

X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale

2025-12-04 · Pei Yang, Hai Ci, Yiren Song, Mike Zheng Shou arxiv

The advancement of embodied AI has unlocked significant potential for intelligent humanoid robots. However, progress in both Vision-Language-Action (VLA) models and world models is severely hampered by the scarcity of la…

UrbanHuRo: A Two-Layer Human-Robot Collaboration Framework for the Joint Optimization of Heterogeneous Urban Services

2026-03-04 · Tonmoy Dey, Lin Jiang, Zheng Dong, Guang Wang arxiv

In the vision of smart cities, technologies are being developed to enhance the efficiency of urban services and improve residents' quality of life. However, most existing research focuses on optimizing individual service…

Reinforcement Learning

ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation

2026-01-31 · Weisheng Dai, Kai Lan, Jianyi Zhou, Bo Zhao 외 arxiv

Vision-Language-Action (VLA) models achieve preliminary generalization through pretraining on large scale robot teleoperation datasets. However, acquiring datasets that comprehensively cover diverse tasks and environment…

Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos

2025-10-24 · Qixiu Li, Yu Deng, Yaobo Liang, Lin Luo 외 arxiv

This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as…

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

2026-06-15 · Hao Li, Ganlong Zhao, Yufei Liu, Haotian Hou 외 arxiv

Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos…