paper-with-me

Papers

OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation

2024-10-15 · Jinhan Li, Yifeng Zhu, Yuqi Xie, Zhenyu Jiang, Mingyo Seo, Georgios Pavlakos, Yuke Zhu

We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for execution. At the heart of our approach is object-aware retargeting, which enables the humanoid robot to mimic the human motions in an RGB-D video while adjusting to different object locations during deployment. OKAMI uses open-world vision models to identify task-relevant objects and retarget the body motions and hand poses separately. Our experiments show that OKAMI achieves strong generalizations across varying visual and spatial conditions, outperforming the state-of-the-art baseline on open-world imitation from observation. Furthermore, OKAMI rollout trajectories are leveraged to train closed-loop visuomotor policies, which achieve an average success rate of 79.2% without the need for labor-intensive teleoperation. More videos can be found on our website https://ut-austin-rpl.github.io/OKAMI/.

📄 PDF Abstract BibTeX arXiv:2410.11792

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generalizable Humanoid Manipulation with 3D Diffusion Policies

2024-10-14 · Yanjie Ze, Zixuan Chen, Wenhao Wang, Tianyi Chen 외

Humanoid robots capable of autonomous operation in diverse environments have long been a goal for roboticists. However, autonomous manipulation by humanoid robots has largely been restricted to one specific scene, primar…

Camera CalibrationPoint Cloud Segmentation

Visual Imitation Enables Contextual Humanoid Control

2025-05-06 · Arthur Allshire, Hongsuk Choi, Junyi Zhang, David McAllister 외

How can we teach humanoids to climb staircases and sit on chairs using the surrounding environment context? Arguably, the simplest way is to just show them-casually capture a human motion video and feed it to humanoids. …

Humanoid Control

RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation

2026-08-04 · Shuliang He, Shuai Wang, Bo Yue, Junchi Teng 외 arxiv

Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive anno…

3D Reconstruction

RHINO: Learning Real-Time Humanoid-Human-Object Interaction from Human Demonstrations

2025-02-18 · Jingxiao Chen, Xinyao Li, Jiahang Cao, Zhengbang Zhu 외

Humanoid robots have shown success in locomotion and manipulation. Despite these basic abilities, humanoids are still required to quickly understand human instructions and react based on human interaction signals to beco…

Human-Object Interaction DetectionObject

OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction

2025-09-30 · Lujie Yang, Xiaoyu Huang, Zhen Wu, Angjoo Kanazawa 외 arxiv

A dominant paradigm for teaching humanoid robots complex skills is to retarget human motions as kinematic references to train reinforcement learning (RL) policies. However, existing retargeting pipelines often struggle w…

Reinforcement LearningData Augmentation