paper-with-me

홈 › Papers

The impacts of known and unknown demonstrator irrationality on reward inference

2021-01-01 · Lawrence Chan, Andrew Critch, Anca Dragan

Algorithms inferring rewards from human behavior typically assume that people are (approximately) rational. In reality, people exhibit a wide array of irrationalities. Motivated by understanding the benefits of modeling these irrationalities, we analyze the effects that demonstrator irrationality has on reward inference. We propose operationalizing several forms of irrationality in the language of MDPs, by altering the Bellman optimality equation, and use this framework to study how these alterations affect inference. We find that incorrectly assuming noisy-rationality for an irrational demonstrator can lead to remarkably poor reward inference accuracy, even in situations where inference with the correct model leads to good inference. This suggests a need to either model irrationalities or find reward inference algorithms that are more robust to misspecification of the demonstrator model. Surprisingly, we find that if we give the learner access to the correct model of the demonstrator's irrationality, these irrationalities can actually help reward inference. In other words, if we could choose between a world where humans were perfectly rational and the current world where humans have systematic biases, the current world might counter-intuitively be preferable for reward inference. We reproduce this effect in several domains. While this finding is mainly conceptual, it is perhaps actionable as well: we might ask human demonstrators for myopic demonstrations instead of optimal ones, as they are more informative for the learner and might be easier for a human to generate.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Inferring Reward Functions from Demonstrators with Unknown Biases

2019-05-01 · ICLR 2019 5 · Rohin Shah, Noah Gundotra, Pieter Abbeel, Anca Dragan

Our goal is to infer reward functions from demonstrations. In order to infer the correct reward function, we must account for the systematic ways in which the demonstrator is suboptimal. Prior work in inverse reinforceme…

Reinforcement Learning

Domain Generalization for Robust Model-Based Offline Reinforcement Learning

2022-11-27 · Alan Clark, Shoaib Ahmed Siddiqui, Robert Kirk, Usman Anwar 외

Existing offline reinforcement learning (RL) algorithms typically assume that training data is either: 1) generated by a known policy, or 2) of entirely unknown origin. We consider multi-demonstrator offline RL, a middle…

Domain GeneralizationOffline RLreinforcement-learningReinforcement Learning+1

Human irrationality: both bad and good for reward inference

2021-11-12 · Lawrence Chan, Andrew Critch, Anca Dragan

Assuming humans are (approximately) rational enables robots to infer reward functions by observing human behavior. But people exhibit a wide array of irrationalities, and our goal with this work is to better understand t…

Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations

2019-07-09 · Daniel S. Brown, Wonjoon Goo, Scott Niekum

The performance of imitation learning is typically upper-bounded by the performance of the demonstrator. While recent empirical results demonstrate that ranked demonstrations allow for better-than-demonstrator performanc…

Imitation Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

VILD: Variational Imitation Learning with Diverse-quality Demonstrations

2019-09-15 · Voot Tangkaratt, Bo Han, Mohammad Emtiyaz Khan, Masashi Sugiyama

The goal of imitation learning (IL) is to learn a good policy from high-quality demonstrations. However, the quality of demonstrations in reality can be diverse, since it is easier and cheaper to collect demonstrations f…

continuous-controlContinuous ControlImitation LearningReinforcement Learning