paper-with-me

홈 › Papers

RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors

2024-12-14 · Fengshuo Bai, Runze Liu, Yali Du, Ying Wen, Yaodong Yang

Evaluating deep reinforcement learning (DRL) agents against targeted behavior attacks is critical for assessing their robustness. These attacks aim to manipulate the victim into specific behaviors that align with the attacker's objectives, often bypassing traditional reward-based defenses. Prior methods have primarily focused on reducing cumulative rewards; however, rewards are typically too generic to capture complex safety requirements effectively. As a result, focusing solely on reward reduction can lead to suboptimal attack strategies, particularly in safety-critical scenarios where more precise behavior manipulation is needed. To address these challenges, we propose RAT, a method designed for universal, targeted behavior attacks. RAT trains an intention policy that is explicitly aligned with human preferences, serving as a precise behavioral target for the adversary. Concurrently, an adversary manipulates the victim's policy to follow this target behavior. To enhance the effectiveness of these attacks, RAT dynamically adjusts the state occupancy measure within the replay buffer, allowing for more controlled and effective behavior manipulation. Our empirical results on robotic simulation tasks demonstrate that RAT outperforms existing adversarial attack algorithms in inducing specific behaviors. Additionally, RAT shows promise in improving agent robustness, leading to more resilient policies. We further validate RAT by guiding Decision Transformer agents to adopt behaviors aligned with human preferences in various MuJoCo tasks, demonstrating its effectiveness across diverse tasks.

📄 PDF Abstract BibTeX arXiv:2412.10713

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial AttackDeep Reinforcement LearningMuJoCo

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
ADOPT Please enter a description about the method here
Adam 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Position-Wise Feed-Forward Layer 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

Query-based Targeted Action-Space Adversarial Policies on Deep Reinforcement Learning Agents

2020-11-13 · Xian Yeow Lee, Yasaman Esfandiari, Kai Liang Tan, Soumik Sarkar

Advances in computing resources have resulted in the increasing complexity of cyber-physical systems (CPS). As the complexity of CPS evolved, the focus has shifted from traditional control methods to deep reinforcement l…

Deep Reinforcement Learningreinforcement-learningReinforcement Learning (RL)Transfer Learning

TrojDRL: Trojan Attacks on Deep Reinforcement Learning Agents

2019-03-01 · Panagiota Kiourti, Kacper Wardega, Susmit Jha, Wenchao Li

Recent work has identified that classification models implemented as neural networks are vulnerable to data-poisoning and Trojan attacks at training time. In this work, we show that these training-time vulnerabilities ex…

Data PoisoningDeep Reinforcement LearningGeneral Classificationreinforcement-learning+2

Targeted Adversarial Attacks on Deep Reinforcement Learning Policies via Model Checking

2022-12-10 · Dennis Gross, Thiago D. Simao, Nils Jansen, Guillermo A. Perez

Deep Reinforcement Learning (RL) agents are susceptible to adversarial noise in their observations that can mislead their policies and decrease their performance. However, an adversary may be interested not only in decre…

Deep Reinforcement Learningreinforcement-learningReinforcement Learning (RL)

A Novel Bifurcation Method for Observation Perturbation Attacks on Reinforcement Learning Agents: Load Altering Attacks on a Cyber Physical Power System

2024-07-06 · Kiernan Broda-Milian, Ranwa Al-Mallah, Hanane Dagdougui

Components of cyber physical systems, which affect real-world processes, are often exposed to the internet. Replacing conventional control methods with Deep Reinforcement Learning (DRL) in energy systems is an active are…

continuous-controlContinuous ControlDeep Reinforcement LearningTime Series Analysis

Neutral Agent-based Adversarial Policy Learning against Deep Reinforcement Learning in Multi-party Open Systems

2025-10-13 · Qizhou Peng, Yang Zheng, Yu Wen, Yanna Wu 외 arxiv

Reinforcement learning (RL) has been an important machine learning paradigm for solving long-horizon sequential decision-making problems under uncertainty. By integrating deep neural networks (DNNs) into the RL framework…

Reinforcement LearningAdversarial AttackAutonomous DrivingStarcraft II