CoinRun: Solving Goal Misgeneralisation
Goal misgeneralisation is a key challenge in AI alignment -- the task of getting powerful Artificial Intelligences to align their goals with human intentions and human morality. In this paper, we show how the ACE (Algorithm for Concept Extrapolation) agent can solve one of the key standard challenges in goal misgeneralisation: the CoinRun challenge. It uses no new reward information in the new environment. This points to how autonomous agents could be trusted to act in human interests, even in novel and critical situations.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Getting By Goal Misgeneralization With a Little Help From a Mentor
While reinforcement learning (RL) agents often perform well during training, they can struggle with distribution shift in real-world deployments. One particularly severe risk of distribution shift is goal misgeneralizati…
Reinforcement Learning (RL)Quantifying Generalization in Reinforcement Learning
In this paper, we investigate the problem of overfitting in deep reinforcement learning. Among the most common benchmarks in RL, it is customary to use the same environments for both training and testing. This practice o…
Data AugmentationDeep Reinforcement LearningL2 Regularizationreinforcement-learning+2The Impossibility of Eliciting Latent Knowledge
Advanced AI systems have extensive knowledge of their environments; in fact, their knowledge may (far) exceed that of their developers or users. Consequently, a desirable property for an AI system is that it is honest --…
Towards Generalizable Reinforcement Learning via Causality-Guided Self-Adaptive Representations
General intelligence requires quick adaption across tasks. While existing reinforcement learning (RL) methods have made progress in generalization, they typically assume only distribution changes between source and targe…
Atari Gamesreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1Instance-based Generalization in Reinforcement Learning
Agents trained via deep reinforcement learning (RL) routinely fail to generalize to unseen environments, even when these share the same underlying dynamics as the training levels. Understanding the generalization propert…
Deep Reinforcement LearningGeneralization Boundsreinforcement-learningReinforcement Learning+1