paper-with-me

홈 › Papers

Penalizing side effects using stepwise relative reachability

2018-06-04 · Victoria Krakovna, Laurent Orseau, Ramana Kumar, Miljan Martic, Shane Legg

How can we design safe reinforcement learning agents that avoid unnecessary disruptions to their environment? We show that current approaches to penalizing side effects can introduce bad incentives, e.g. to prevent any irreversible changes in the environment, including the actions of other agents. To isolate the source of such undesirable incentives, we break down side effects penalties into two components: a baseline state and a measure of deviation from this baseline state. We argue that some of these incentives arise from the choice of baseline, and others arise from the choice of deviation measure. We introduce a new variant of the stepwise inaction baseline and a new deviation measure based on relative reachability of states. The combination of these design choices avoids the given undesirable incentives, while simpler baselines and the unreachability measure fail. We demonstrate this empirically by comparing different combinations of baseline and deviation measure choices on a set of gridworld experiments designed to illustrate possible bad incentives.

📄 PDF Abstract BibTeX arXiv:1806.01186

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningSafe Reinforcement Learning

Similar Papers 제목 키워드 기반

Avoiding Side Effects in Complex Environments

2020-06-11 · NeurIPS 2020 12 · Alexander Matt Turner, Neale Ratzlaff, Prasad Tadepalli

Reward function specification can be difficult. Rewarding the agent for making a widget may be easy, but penalizing the multitude of possible negative side effects is hard. In toy environments, Attainable Utility Preserv…

Go Beyond Imagination: Maximizing Episodic Reachability with World Models

2023-08-25 · Yao Fu, Run Peng, Honglak Lee

Efficient exploration is a challenging topic in reinforcement learning, especially for sparse reward tasks. To deal with the reward sparsity, people commonly apply intrinsic rewards to motivate agents to explore the stat…

Efficient Exploration

Quantifying Availability and Discovery in Recommender Systems via Stochastic Reachability

2021-06-30 · Mihaela Curmei, Sarah Dean, Benjamin Recht

In this work, we consider how preference models in interactive recommendation systems determine the availability of content and users' opportunities for discovery. We propose an evaluation procedure based on stochastic r…

Interactive RecommendationRecommendation Systems

Avoiding Side Effects By Considering Future Tasks

2020-10-15 · NeurIPS 2020 12 · Victoria Krakovna, Laurent Orseau, Richard Ngo, Miljan Martic 외

Designing reward functions is difficult: the designer has to specify what to do (what it means to complete the task) as well as what not to do (side effects that should be avoided while completing the task). To alleviate…

Skill Reuse as Compression in Agentic RL

2026-05-29 · Zhikun Xu, Yu Feng, Jacob Dineen, Taiwei Shi 외 arxiv

Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generalize better when their successful trajectories are structurally compress…

Reinforcement Learning