paper-with-me

Papers

Generalizable task representation learning from human demonstration videos: a geometric approach

2022-02-28 · Jun Jin, Martin Jagersand

We study the problem of generalizable task learning from human demonstration videos without extra training on the robot or pre-recorded robot motions. Given a set of human demonstration videos showing a task with different objects/tools (categorical objects), we aim to learn a representation of visual observation that generalizes to categorical objects and enables efficient controller design. We propose to introduce a geometric task structure to the representation learning problem that geometrically encodes the task specification from human demonstration videos, and that enables generalization by building task specification correspondence between categorical objects. Specifically, we propose CoVGS-IL, which uses a graph-structured task function to learn task representations under structural constraints. Our method enables task generalization by selecting geometric features from different objects whose inner connection relationships define the same task in geometric constraints. The learned task representation is then transferred to a robot controller using uncalibrated visual servoing (UVS); thus, the need for extra robot training or pre-recorded robot motions is removed.

📄 PDF Abstract BibTeX arXiv:2202.13604

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

2024-03-06 · Yanjie Ze, Gu Zhang, Kangning Zhang, Chenyuan Hu 외

Imitation learning provides an efficient way to teach robots dexterous skills; however, learning complex skills robustly and generalizablely usually consumes large amounts of human demonstrations. To tackle this challeng…

Imitation LearningRobot Manipulation

Mitty: Diffusion-based Human-to-Robot Video Generation

2025-12-19 · Yiren Song, Cheng Liu, Weijia Mao, Mike Zheng Shou arxiv

Learning directly from human demonstration videos is a key milestone toward scalable and generalizable robot learning. Yet existing methods rely on intermediate representations such as keypoints or trajectories, introduc…

Video Generation

Giving Robots a Hand: Learning Generalizable Manipulation with Eye-in-Hand Human Video Demonstrations

2023-07-12 · Moo Jin Kim, Jiajun Wu, Chelsea Finn

Eye-in-hand cameras have shown promise in enabling greater sample efficiency and generalization in vision-based robotic manipulation. However, for robotic imitation, it is still expensive to have a human teleoperator col…

Domain Adaptation

Parse-Augment-Distill: Learning Generalizable Bimanual Visuomotor Policies from Single Human Video

2025-09-24 · Georgios Tziafas, Jiayun Zhang, Hamidreza Kasaei arxiv

Learning visuomotor policies from expert demonstrations is an important frontier in modern robotics research, however, most popular methods require copious efforts for collecting teleoperation data and struggle to genera…

Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

2024-05-02 · Homanga Bharadhwaj, Roozbeh Mottaghi, Abhinav Gupta, Shubham Tulsiani

We seek to learn a generalizable goal-conditioned policy that enables zero-shot robot manipulation: interacting with unseen objects in novel scenes without test-time adaptation. While typical approaches rely on a large a…

Robot ManipulationTest-time Adaptation