paper-with-me

홈 › Papers

X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale

2025-12-04 · Pei Yang, Hai Ci, Yiren Song, Mike Zheng Shou arxiv

The advancement of embodied AI has unlocked significant potential for intelligent humanoid robots. However, progress in both Vision-Language-Action (VLA) models and world models is severely hampered by the scarcity of large-scale, diverse training data. A promising solution is to "robotize" web-scale human videos, which has been proven effective for policy training. However, these solutions mainly "overlay" robot arms to egocentric videos, which cannot handle complex full-body motions and scene occlusions in third-person videos, making them unsuitable for robotizing humans. To bridge this gap, we introduce X-Humanoid, a generative video editing approach that adapts the powerful Wan 2.2 model into a video-to-video structure and finetunes it for the human-to-humanoid translation task. This finetuning requires paired human-humanoid videos, so we designed a scalable data creation pipeline, turning community assets into 17+ hours of paired synthetic videos using Unreal Engine. We then apply our trained model to 60 hours of the Ego-Exo4D videos, generating and releasing a new large-scale dataset of over 3.6 million "robotized" humanoid video frames. Quantitative analysis and user studies confirm our method's superiority over existing baselines: 69% of users rated it best for motion consistency, and 62.1% for embodiment correctness.

📄 PDF Abstract BibTeX arXiv:2512.04537

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos

2025-09-20 · Haoyang Weng, Yitang Li, Nikhil Sobanbabu, Zihan Wang 외 arxiv

Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for Interaction), a simple and general framewo…

Reinforcement Learning

Learning from Massive Human Videos for Universal Humanoid Pose Control

2024-12-18 · Jiageng Mao, Siheng Zhao, Siqi Song, Tianheng Shi 외

Scalable learning of humanoid robots is crucial for their deployment in real-world applications. While traditional approaches primarily rely on reinforcement learning or teleoperation to achieve whole-body control, they …

Caption GenerationHumanoid Controlmotion retargeting

Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments

2026-06-30 · Xiaopeng Lin, Ruoqi Yang, Shijie Lian, Zhaolong Shen 외 arxiv

Vision-language-action (VLA) models across robot embodiments require high-quality observation--action supervision to learn deployable action distributions, yet scaling such robot data remains difficult, especially for hi…

Humanoid Motion Scripting with Postural Synergies

2025-08-17 · Rhea Malhotra, William Chong, Catie Cuan, Oussama Khatib arxiv

Generating sequences of human-like motions for humanoid robots presents challenges in collecting and analyzing reference human motions, synthesizing new motions based on these reference motions, and mapping the generated…

Unveiling the Impact of Data and Model Scaling on High-Level Control for Humanoid Robots

2025-11-12 · Yuxi Wei, Zirui Wang, Kangning Yin, Yue Hu 외 arxiv

Data scaling has long remained a critical bottleneck in robot learning. For humanoid robots, human videos and motion data are abundant and widely available, offering a free and large-scale data source. Besides, the seman…