paper-with-me

홈 › Papers

H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation

2025-07-31 · Hongzhe Bi, Lingxuan Wu, Tianwei Lin, Hengkai Tan, Zhizhong Su, Hang Su, Jun Zhu arxiv

Imitation learning for robotic manipulation faces a fundamental challenge: the scarcity of large-scale, high-quality robot demonstration data. Recent robotic foundation models often pre-train on cross-embodiment robot datasets to increase data scale, while they face significant limitations as the diverse morphologies and action spaces across different robot embodiments make unified training challenging. In this paper, we present H-RDT (Human to Robotics Diffusion Transformer), a novel approach that leverages human manipulation data to enhance robot manipulation capabilities. Our key insight is that large-scale egocentric human manipulation videos with paired 3D hand pose annotations provide rich behavioral priors that capture natural manipulation strategies and can benefit robotic policy learning. We introduce a two-stage training paradigm: (1) pre-training on large-scale egocentric human manipulation data, and (2) cross-embodiment fine-tuning on robot-specific data with modular action encoders and decoders. Built on a diffusion transformer architecture with 2B parameters, H-RDT uses flow matching to model complex action distributions. Extensive evaluations encompassing both simulation and real-world experiments, single-task and multitask scenarios, as well as few-shot learning and robustness assessments, demonstrate that H-RDT outperforms training from scratch and existing state-of-the-art methods, including Pi0 and RDT, achieving significant improvements of 13.9% and 40.5% over training from scratch in simulation and real-world experiments, respectively. The results validate our core hypothesis that human manipulation data can serve as a powerful foundation for learning bimanual robotic manipulation policies.

📄 PDF Abstract BibTeX arXiv:2507.23523

Code (0)

등록된 구현이 없습니다.

Tasks

Robot ManipulationFew-Shot Learning

Similar Papers 제목 키워드 기반

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

2025-11-21 · Shihan Wu, Xuecheng Liu, Shaoxuan Xie, Pengwei Wang 외 arxiv

Despite the critical role of bimanual manipulation in endowing robots with human-like dexterity, large-scale and diverse datasets remain scarce due to the significant hardware heterogeneity across bimanual robotic platfo…

BiCoord: A Bimanual Manipulation Benchmark towards Long-Horizon Spatial-Temporal Coordination

2026-04-07 · Xingyu Peng, Chen Gao, Liankai Jin, Annan Li 외 arxiv

Bimanual manipulation, i.e., the coordinated use of two robotic arms to complete tasks, is essential for achieving human-level dexterity in robotics. Recent simulation benchmarks, e.g., RoboTwin and RLBench2, have advanc…

ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning

2025-03-27 · CVPR 2025 1 · Kailin Li, Puhao Li, Tengyu Liu, Yuyang Li 외

Human hands play a central role in interacting, motivating increasing research in dexterous robotic manipulation. Data-driven embodied AI algorithms demand precise, large-scale, human-like manipulation sequences, which a…

ScrewMimic: Bimanual Imitation from Human Videos with Screw Space Projection

2024-05-06 · Arpit Bahety, Priyanka Mandikal, Ben Abbatematteo, Roberto Martín-Martín

Bimanual manipulation is a longstanding challenge in robotics due to the large number of degrees of freedom and the strict spatial and temporal synchronization required to generate meaningful behavior. Humans learn biman…

DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos

2026-02-10 · Juncheng Mu, Sizhe Yang, Yiming Bao, Hojin Bae 외 arxiv

Data scarcity fundamentally limits the generalization of bimanual dexterous manipulation, as real-world data collection for dexterous hands is expensive and labor-intensive. Human manipulation videos, as a direct carrier…

Data AugmentationVideo Generation