paper-with-me

홈 › Papers

Avoiding Side Effects in Complex Environments

2020-06-11 · NeurIPS 2020 12 · Alexander Matt Turner, Neale Ratzlaff, Prasad Tadepalli

Reward function specification can be difficult. Rewarding the agent for making a widget may be easy, but penalizing the multitude of possible negative side effects is hard. In toy environments, Attainable Utility Preservation (AUP) avoided side effects by penalizing shifts in the ability to achieve randomly generated goals. We scale this approach to large, randomly generated environments based on Conway's Game of Life. By preserving optimal value for a single randomly generated reward function, AUP incurs modest overhead while leading the agent to complete the specified task and avoid many side effects. Videos and code are available at https://avoiding-side-effects.github.io/.

📄 PDF Abstract BibTeX arXiv:2006.06547

Code (2)

aseembits93/avoiding-side-effects pytorch
neale/avoiding-side-effects pytorch

Similar Papers 제목 키워드 기반

Avoiding Side Effects By Considering Future Tasks

2020-10-15 · NeurIPS 2020 12 · Victoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic 외

Designing reward functions is difficult: the designer has to specify what to do (what it means to complete the task) as well as what not to do (side effects that should be avoided while completing the task). To alleviate…

Avoiding Negative Side Effects due to Incomplete Knowledge of AI Systems

2020-08-24 · Sandhya Saisubramanian, Shlomo Zilberstein, Ece Kamar

Autonomous agents acting in the real-world often operate based on models that ignore certain aspects of the environment. The incompleteness of any given model -- handcrafted or machine acquired -- is inevitable due to pr…

AI Safety Gridworlds

2017-11-27 · Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A. Ortega 외

We present a suite of reinforcement learning environments illustrating various safety properties of intelligent agents. These problems include safe interruptibility, avoiding side effects, absent supervisor, reward gamin…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

SafeLife 1.0: Exploring Side Effects in Complex Environments

2019-12-03 · Carroll L. Wainwright, Peter Eckersley

We present SafeLife, a publicly available reinforcement learning environment that tests the safety of reinforcement learning agents. It contains complex, dynamic, tunable, procedurally generated levels with many opportun…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

On Avoiding Power-Seeking by Artificial Intelligence

2022-06-23 · Alexander Matt Turner

We do not know how to align a very intelligent AI agent's behavior with human interests. I investigate whether -- absent a full solution to this AI alignment problem -- we can build smart AI agents which have limited imp…

Decision Making