paper-with-me

Papers

EmbodiSwap for Zero-Shot Robot Imitation Learning

2025-10-04 · Eadom Dessalene, Pavan Mantripragada, Michael Maynord, Yiannis Aloimonos arxiv

We introduce EmbodiSwap - a method for producing photorealistic synthetic robot overlays over human video. We employ EmbodiSwap for zero-shot imitation learning, bridging the embodiment gap between in-the-wild ego-centric human video and a target robot embodiment. We train a closed-loop robot manipulation policy over the data produced by EmbodiSwap. We make novel use of V-JEPA as a visual backbone, repurposing V-JEPA from the domain of video understanding to imitation learning over synthetic robot videos. Adoption of V-JEPA outperforms alternative vision backbones more conventionally used within robotics. In real-world tests, our zero-shot trained V-JEPA model achieves an $82\%$ success rate, outperforming a few-shot trained $π_0$ network as well as $π_0$ trained over data produced by EmbodiSwap. We release (i) code for generating the synthetic robot overlays which takes as input human videos and an arbitrary robot URDF and generates a robot dataset, (ii) the robot dataset we synthesize over EPIC-Kitchens, HOI4D and Ego4D, and (iii) model checkpoints and inference code, to facilitate reproducible research and broader adoption.

📄 PDF Abstract BibTeX arXiv:2510.03706

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

Zero-shot Imitation Learning from Demonstrations for Legged Robot Visual Navigation

2019-09-27 · Xinlei Pan, Tingnan Zhang, Brian Ichter, Aleksandra Faust 외

Imitation learning is a popular approach for training visual navigation policies. However, collecting expert demonstrations for legged robots is challenging as these robots can be hard to control, move slowly, and cannot…

DisentanglementImitation LearningVisual Navigation

Zero-Shot Robot Manipulation from Passive Human Videos

2023-02-03 · Homanga Bharadhwaj, Abhinav Gupta, Shubham Tulsiani, Vikash Kumar

Can we learn robot manipulation for everyday tasks, only by watching videos of humans doing arbitrary tasks in different unstructured settings? Unlike widely adopted strategies of learning task-specific behaviors or dire…

Robot Manipulation

Zero-Shot Visual Imitation

2018-04-23 · ICLR 2018 1 · Deepak Pathak, Parsa Mahmoudieh, Guanghao Luo, Pulkit Agrawal 외

The current dominant paradigm for imitation learning relies on strong supervision of expert actions to learn both 'what' and 'how' to imitate. We pursue an alternative paradigm wherein an agent first explores the world w…

Imitation Learning

EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment

2026-06-11 · Shihefeng Wang, Kangchen Lv, Mingrui Yu, Xiang Li arxiv

Scalable robot imitation learning relies on large-scale heterogeneous data from diverse robots or body-free data, making Cartesian end-effector actions a key interface for embodiment-agnostic policy learning. However, en…

Collision Avoidance

Reliable Semantic Understanding for Real World Zero-shot Object Goal Navigation

2024-10-29 · Halil Utku Unlu, Shuaihang Yuan, Congcong Wen, Hao Huang 외

We introduce an innovative approach to advancing semantic understanding in zero-shot object goal navigation (ZS-OGN), enhancing the autonomy of robots in unfamiliar environments. Traditional reliance on labeled data has …

Decision MakingLanguage ModelingLanguage ModellingObject