paper-with-me

홈 › Papers

Reinforcement Learning with a Corrupted Reward Channel

2017-05-23 · Tom Everitt, Victoria Krakovna, Laurent Orseau, Marcus Hutter, Shane Legg

No real-world reward function is perfect. Sensory errors and software bugs may result in RL agents observing higher (or lower) rewards than they should. For example, a reinforcement learning agent may prefer states where a sensory error gives it the maximum reward, but where the true reward is actually small. We formalise this problem as a generalised Markov Decision Problem called Corrupt Reward MDP. Traditional RL methods fare poorly in CRMDPs, even under strong simplifying assumptions and when trying to compensate for the possibly corrupt rewards. Two ways around the problem are investigated. First, by giving the agent richer data, such as in inverse reinforcement learning and semi-supervised reinforcement learning, reward corruption stemming from systematic sensory errors may sometimes be completely managed. Second, by using randomisation to blunt the agent's optimisation, reward corruption can be partially managed under some assumptions.

📄 PDF Abstract BibTeX arXiv:1705.08417

Code (1)

jvmancuso/safe-grid-agents tf

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Reinforcement Learning with Perturbed Rewards

2018-10-02 · ICLR 2019 5 · Jingkang Wang, Yang Liu, Bo Li

Recent studies have shown that reinforcement learning (RL) models are vulnerable in various noisy scenarios. For instance, the observed reward channel is often subject to noise in practice (e.g., when rewards are collect…

Atari GamesDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1

Reward Estimation for Variance Reduction in Deep Reinforcement Learning

2018-05-09 · Joshua Romoff, Peter Henderson, Alexandre Piché, Vincent Francois-Lavet 외

Reinforcement Learning (RL) agents require the specification of a reward signal for learning behaviours. However, introduction of corrupt or stochastic rewards can yield high variance in learning. Such corruption may be …

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Distributional Reward Decomposition for Reinforcement Learning

2019-11-06 · NeurIPS 2019 12 · Zichuan Lin, Li Zhao, Derek Yang, Tao Qin 외

Many reinforcement learning (RL) tasks have specific properties that can be leveraged to modify existing RL algorithms to adapt to those tasks and further improve performance, and a general class of such properties is th…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Improved Corruption Robust Algorithms for Episodic Reinforcement Learning

2021-02-13 · Yifang Chen, Simon S. Du, Kevin Jamieson

We study episodic reinforcement learning under unknown adversarial corruptions in both the rewards and the transition probabilities of the underlying system. We propose new algorithms which, compared to the existing resu…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Robust Asynchronous Q-Learning under Reward and State Corruption via Batching

2026-07-23 · Sreejeet Maity, Aritra Mitra arxiv

Motivated by reinforcement learning in harsh environments, we consider the problem of learning an optimal policy subject to adversarially corrupted feedback. Specifically, at each time-step, an adversary can perturb both…

Reinforcement Learning