paper-with-me

Papers

Incorrigibility in the CIRL Framework

2017-09-19 · Ryan Carey

A value learning system has incentives to follow shutdown instructions, assuming the shutdown instruction provides information (in the technical sense) about which actions lead to valuable outcomes. However, this assumption is not robust to model mis-specification (e.g., in the case of programmer errors). We demonstrate this by presenting some Supervised POMDP scenarios in which errors in the parameterized reward function remove the incentive to follow shutdown commands. These difficulties parallel those discussed by Soares et al. (2015) in their paper on corrigibility. We argue that it is important to consider systems that follow shutdown commands under some weaker set of assumptions (e.g., that one small verified module is correctly implemented; as opposed to an entire prior probability distribution and/or parameterized reward function). We discuss some difficulties with simple ways to attempt to attain these sorts of guarantees in a value learning framework.

📄 PDF Abstract BibTeX arXiv:1709.06275

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Control-Informed Reinforcement Learning for Chemical Processes

2024-08-24 · Maximilian Bloor, Akhil Ahmed, Niki Kotecha, Mehmet Mercangöz 외

This work proposes a control-informed reinforcement learning (CIRL) framework that integrates proportional-integral-derivative (PID) control components into the architecture of deep reinforcement learning (RL) policies. …

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Cooperative Inverse Reinforcement Learning

2016-06-09 · NeurIPS 2016 12 · Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, Stuart Russell

For an autonomous system to be helpful to humans and to pose no unwarranted risks, it needs to align its values with those of the humans in its environment in such a way that its actions contribute to the maximization of…

Active Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

An Efficient, Generalized Bellman Update For Cooperative Inverse Reinforcement Learning

2018-06-11 · ICML 2018 7 · Dhruv Malik, Malayandi Palaniappan, Jaime F. Fisac, Dylan Hadfield-Menell 외

Our goal is for AI systems to correctly identify and act according to their human user's objectives. Cooperative Inverse Reinforcement Learning (CIRL) formalizes this value alignment problem as a two-player game between …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Non-Cooperative Inverse Reinforcement Learning

2019-11-03 · NeurIPS 2019 12 · Xiangyuan Zhang, Kaiqing Zhang, Erik Miehling, Tamer Başar

Making decisions in the presence of a strategic opponent requires one to take into account the opponent's ability to actively mask its intended objective. To describe such strategic situations, we introduce the non-coope…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Interaction-limited Inverse Reinforcement Learning

2020-07-01 · Martin Troussard, Emmanuel Pignat, Parameswaran Kamalaruban, Sylvain Calinon 외

This paper proposes an inverse reinforcement learning (IRL) framework to accelerate learning when the learner-teacher \textit{interaction} is \textit{limited} during training. Our setting is motivated by the realistic sc…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)