paper-with-me

홈 › Papers

Inferring Reward Functions from Demonstrators with Unknown Biases

2019-05-01 · ICLR 2019 5 · Rohin Shah, Noah Gundotra, Pieter Abbeel, Anca Dragan

Our goal is to infer reward functions from demonstrations. In order to infer the correct reward function, we must account for the systematic ways in which the demonstrator is suboptimal. Prior work in inverse reinforcement learning can account for specific, known biases, but cannot handle demonstrators with unknown biases. In this work, we explore the idea of learning the demonstrator's planning algorithm (including their unknown biases), along with their reward function. What makes this challenging is that any demonstration could be explained either by positing a term in the reward function, or by positing a particular systematic bias. We explore what assumptions are sufficient for avoiding this impossibility result: either access to tasks with known rewards which enable estimating the planner separately, or that the demonstrator is sufficiently close to optimal that this can serve as a regularizer. In our exploration with synthetic models of human biases, we find that it is possible to adapt to different biases and perform better than assuming a fixed model of the demonstrator, such as Boltzmann rationality.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

The impacts of known and unknown demonstrator irrationality on reward inference

2021-01-01 · Lawrence Chan, Andrew Critch, Anca Dragan

Algorithms inferring rewards from human behavior typically assume that people are (approximately) rational. In reality, people exhibit a wide array of irrationalities. Motivated by understanding the benefits of modeling …

Joint Goal and Strategy Inference across Heterogeneous Demonstrators via Reward Network Distillation

2020-01-02 · Letian Chen, Rohan Paleja, Muyleng Ghuy, Matthew Gombolay

Reinforcement learning (RL) has achieved tremendous success as a general framework for learning how to make decisions. However, this success relies on the interactive hand-tuning of a reward function by RL experts. On th…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Domain Generalization for Robust Model-Based Offline Reinforcement Learning

2022-11-27 · Alan Clark, Shoaib Ahmed Siddiqui, Robert Kirk, Usman Anwar 외

Existing offline reinforcement learning (RL) algorithms typically assume that training data is either: 1) generated by a known policy, or 2) of entirely unknown origin. We consider multi-demonstrator offline RL, a middle…

Domain GeneralizationOffline RLreinforcement-learningReinforcement Learning+1

Deep Adaptive Multi-Intention Inverse Reinforcement Learning

2021-07-14 · Ariyan Bighashdel, Panagiotis Meletis, Pavol Jancura, Gijs Dubbelman

This paper presents a deep Inverse Reinforcement Learning (IRL) framework that can learn an a priori unknown number of nonlinear reward functions from unlabeled experts' demonstrations. For this purpose, we employ the to…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

VILD: Variational Imitation Learning with Diverse-quality Demonstrations

2019-09-15 · Voot Tangkaratt, Bo Han, Mohammad Emtiyaz Khan, Masashi Sugiyama

The goal of imitation learning (IL) is to learn a good policy from high-quality demonstrations. However, the quality of demonstrations in reality can be diverse, since it is easier and cheaper to collect demonstrations f…

continuous-controlContinuous ControlImitation LearningReinforcement Learning