Execute Order 66: Targeted Data Poisoning for Reinforcement Learning
Data poisoning for reinforcement learning has historically focused on general performance degradation, and targeted attacks have been successful via perturbations that involve control of the victim's policy and rewards. We introduce an insidious poisoning attack for reinforcement learning which causes agent misbehavior only at specific target states - all while minimally modifying a small fraction of training observations without assuming any control over policy or reward. We accomplish this by adapting a recent technique, gradient alignment, to reinforcement learning. We test our method and demonstrate success in two Atari games of varying difficulty.
Code (0)
등록된 구현이 없습니다.
Tasks
Atari GamesData Poisoningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Understanding the Limits of Poisoning Attacks in Episodic Reinforcement Learning
To understand the security threats to reinforcement learning (RL) algorithms, this paper studies poisoning attacks to manipulate \emph{any} order-optimal learning algorithm towards a targeted policy in episodic RL and ex…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks
Contrastive Language-Image Pre-training (CLIP) on large image-caption datasets has achieved remarkable success in zero-shot classification and enabled transferability to new domains. However, CLIP is extremely more vulne…
Contrastive LearningData Poisoningzero-shot-classificationZero-Shot LearningBlack-Box Targeted Reward Poisoning Attack Against Online Deep Reinforcement Learning
We propose the first black-box targeted attack against online deep reinforcement learning through reward poisoning during training time. Our attack is applicable to general environments with unknown dynamics learned by u…
Deep Reinforcement Learningreinforcement-learningTrojDRL: Trojan Attacks on Deep Reinforcement Learning Agents
Recent work has identified that classification models implemented as neural networks are vulnerable to data-poisoning and Trojan attacks at training time. In this work, we show that these training-time vulnerabilities ex…
Data PoisoningDeep Reinforcement LearningGeneral Classificationreinforcement-learning+2BadRL: Sparse Targeted Backdoor Attack Against Reinforcement Learning
Backdoor attacks in reinforcement learning (RL) have previously employed intense attack strategies to ensure attack success. However, these methods suffer from high attack costs and increased detectability. In this work,…
Backdoor Attackreinforcement-learningReinforcement LearningReinforcement Learning (RL)