paper-with-me

홈 › Papers

Disturbing Reinforcement Learning Agents with Corrupted Rewards

2021-02-12 · Rubén Majadas, Javier García, Fernando Fernández

Reinforcement Learning (RL) algorithms have led to recent successes in solving complex games, such as Atari or Starcraft, and to a huge impact in real-world applications, such as cybersecurity or autonomous driving. In the side of the drawbacks, recent works have shown how the performance of RL algorithms decreases under the influence of soft changes in the reward function. However, little work has been done about how sensitive these disturbances are depending on the aggressiveness of the attack and the learning exploration strategy. In this paper, we propose to fill this gap in the literature analyzing the effects of different attack strategies based on reward perturbations, and studying the effect in the learner depending on its exploration strategy. In order to explain all the behaviors, we choose a sub-class of MDPs: episodic, stochastic goal-only-rewards MDPs, and in particular, an intelligible grid domain as a benchmark. In this domain, we demonstrate that smoothly crafting adversarial rewards are able to mislead the learner, and that using low exploration probability values, the policy learned is more robust to corrupt rewards. Finally, in the proposed learning scenario, a counterintuitive result arises: attacking at each learning episode is the lowest cost attack strategy.

📄 PDF Abstract BibTeX arXiv:2102.06587

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Drivingreinforcement-learningReinforcement LearningReinforcement Learning (RL)Starcraft

Similar Papers 제목 키워드 기반

Reward Estimation for Variance Reduction in Deep Reinforcement Learning

2018-05-09 · Joshua Romoff, Peter Henderson, Alexandre Piché, Vincent Francois-Lavet 외

Reinforcement Learning (RL) agents require the specification of a reward signal for learning behaviours. However, introduction of corrupt or stochastic rewards can yield high variance in learning. Such corruption may be …

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

The Distributional Reward Critic Framework for Reinforcement Learning Under Perturbed Rewards

2024-01-11 · Xi Chen, Zhihui Zhu, Andrew Perrault

The reward signal plays a central role in defining the desired behaviors of agents in reinforcement learning (RL). Rewards collected from realistic environments could be perturbed, corrupted, or noisy due to an adversary…

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning+1

Reinforcement Learning with a Corrupted Reward Channel

2017-05-23 · Tom Everitt, Victoria Krakovna, Laurent Orseau, Marcus Hutter 외

No real-world reward function is perfect. Sensory errors and software bugs may result in RL agents observing higher (or lower) rewards than they should. For example, a reinforcement learning agent may prefer states where…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Reinforcement Learning with Perturbed Rewards

2018-10-02 · ICLR 2019 5 · Jingkang Wang, Yang Liu, Bo Li

Recent studies have shown that reinforcement learning (RL) models are vulnerable in various noisy scenarios. For instance, the observed reward channel is often subject to noise in practice (e.g., when rewards are collect…

Atari GamesDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1

Uncertainty-based Offline Variational Bayesian Reinforcement Learning for Robustness under Diverse Data Corruptions

2024-11-01 · Rui Yang, Jie Wang, Guoping Wu, Bin Li

Real-world offline datasets are often subject to data corruptions (such as noise or adversarial attacks) due to sensor failures or malicious attacks. Despite advances in robust offline reinforcement learning (RL), existi…

Bayesian InferenceOffline RLReinforcement Learning (RL)