paper-with-me

홈 › Papers

Learning Reward Functions for Robotic Manipulation by Observing Humans

2022-11-16 · Minttu Alakuijala, Gabriel Dulac-Arnold, Julien Mairal, Jean Ponce, Cordelia Schmid

Observing a human demonstrator manipulate objects provides a rich, scalable and inexpensive source of data for learning robotic policies. However, transferring skills from human videos to a robotic manipulator poses several challenges, not least a difference in action and observation spaces. In this work, we use unlabeled videos of humans solving a wide range of manipulation tasks to learn a task-agnostic reward function for robotic manipulation policies. Thanks to the diversity of this training data, the learned reward function sufficiently generalizes to image observations from a previously unseen robot embodiment and environment to provide a meaningful prior for directed exploration in reinforcement learning. We propose two methods for scoring states relative to a goal image: through direct temporal regression, and through distances in an embedding space obtained with time-contrastive learning. By conditioning the function on a goal image, we are able to reuse one model across a variety of tasks. Unlike prior work on leveraging human videos to teach robots, our method, Human Offline Learned Distances (HOLD) requires neither a priori data from the robot environment, nor a set of task-specific human demonstrations, nor a predefined notion of correspondence across morphologies, yet it is able to accelerate training of several manipulation tasks on a simulated robot arm compared to using only a sparse reward obtained from task completion.

📄 PDF Abstract BibTeX arXiv:2211.09019

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

Supervised Reward Inference

2025-02-25 · Will Schwarzer, Jordan Schneider, Philip S. Thomas, Scott Niekum

Existing approaches to reward inference from behavior typically assume that humans provide demonstrations according to specific models of behavior. However, humans often indicate their goals through a wide range of behav…

Learning Generalizable Robotic Reward Functions from "In-The-Wild" Human Videos

2021-03-09 · ICLR Workshop SSL-RL 2021 5 · Anonymous

We are motivated by the goal of generalist robotic agents that can complete a wide range of tasks across many environments. Critical to this is the robot’s ability to acquire some metric of task success or reward, which …

Model Predictive Control

LARG, Language-based Automatic Reward and Goal Generation

2023-06-19 · Julien Perez, Denys Proux, Claude Roux, Michael Niemaz

Goal-conditioned and Multi-Task Reinforcement Learning (GCRL and MTRL) address numerous problems related to robot learning, including locomotion, navigation, and manipulation scenarios. Recent works focusing on language-…

reinforcement-learningReinforcement Learning

MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models

2026-01-28 · Xunlan Zhou, Xuanlin Chen, Shaowei Zhang, ShengHua Wan 외 arxiv

Designing dense reward functions is pivotal for efficient robotic Reinforcement Learning (RL). However, most dense rewards rely on manual engineering, which fundamentally limits the scalability and automation of reinforc…

Reinforcement Learning

Learning Predictive Models From Observation and Interaction

2019-12-30 · ECCV 2020 8 · Karl Schmeckpeper, Annie Xie, Oleh Rybkin, Stephen Tian 외

Learning predictive models from interaction with the world allows an agent, such as a robot, to learn about how the world works, and then use this learned model to plan coordinated sequences of actions to bring about des…