Incorrigibility in the CIRL Framework
A value learning system has incentives to follow shutdown instructions, assuming the shutdown instruction provides information (in the technical sense) about which actions lead to valuable outcomes. However, this assumption is not robust to model mis-specification (e.g., in the case of programmer errors). We demonstrate this by presenting some Supervised POMDP scenarios in which errors in the parameterized reward function remove the incentive to follow shutdown commands. These difficulties parallel those discussed by Soares et al. (2015) in their paper on corrigibility. We argue that it is important to consider systems that follow shutdown commands under some weaker set of assumptions (e.g., that one small verified module is correctly implemented; as opposed to an entire prior probability distribution and/or parameterized reward function). We discuss some difficulties with simple ways to attempt to attain these sorts of guarantees in a value learning framework.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Control-Informed Reinforcement Learning for Chemical Processes
This work proposes a control-informed reinforcement learning (CIRL) framework that integrates proportional-integral-derivative (PID) control components into the architecture of deep reinforcement learning (RL) policies. …
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Cooperative Inverse Reinforcement Learning
For an autonomous system to be helpful to humans and to pose no unwarranted risks, it needs to align its values with those of the humans in its environment in such a way that its actions contribute to the maximization of…
Active Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)An Efficient, Generalized Bellman Update For Cooperative Inverse Reinforcement Learning
Our goal is for AI systems to correctly identify and act according to their human user's objectives. Cooperative Inverse Reinforcement Learning (CIRL) formalizes this value alignment problem as a two-player game between …
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Non-Cooperative Inverse Reinforcement Learning
Making decisions in the presence of a strategic opponent requires one to take into account the opponent's ability to actively mask its intended objective. To describe such strategic situations, we introduce the non-coope…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Interaction-limited Inverse Reinforcement Learning
This paper proposes an inverse reinforcement learning (IRL) framework to accelerate learning when the learner-teacher \textit{interaction} is \textit{limited} during training. Our setting is motivated by the realistic sc…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)