Policy Disruption in Reinforcement Learning:Adversarial Attack with Large Language Models and Critical State Identification
Reinforcement learning (RL) has achieved remarkable success in fields like robotics and autonomous driving, but adversarial attacks designed to mislead RL systems remain challenging. Existing approaches often rely on modifying the environment or policy, limiting their practicality. This paper proposes an adversarial attack method in which existing agents in the environment guide the target policy to output suboptimal actions without altering the environment. We propose a reward iteration optimization framework that leverages large language models (LLMs) to generate adversarial rewards explicitly tailored to the vulnerabilities of the target agent, thereby enhancing the effectiveness of inducing the target agent toward suboptimal decision-making. Additionally, a critical state identification algorithm is designed to pinpoint the target agent's most vulnerable states, where suboptimal behavior from the victim leads to significant degradation in overall performance. Experimental results in diverse environments demonstrate the superiority of our method over existing approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
Reinforcement LearningAdversarial AttackAutonomous DrivingSimilar Papers 제목 키워드 기반
Reinforcement Learning Disrupts Gradient-Based Adversarial Optimization
Gradient-based adversarial attacks remain a dominant threat to deep neural networks (DNNs), as they exploit gradient information to efficiently optimize adversarial perturbations. To address this, we investigate whether …
Reinforcement LearningSparse Adversarial Attack in Multi-agent Reinforcement Learning
Cooperative multi-agent reinforcement learning (cMARL) has many real applications, but the policy trained by existing cMARL algorithms is not robust enough when deployed. There exist also many methods about adversarial a…
Adversarial AttackMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+1Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement Learning
Cooperation is central to multi-agent reinforcement learning (MARL), yet learned coordination can be fragile when external perturbations disrupt inter-agent interactions. Prior robust MARL methods have primarily consider…
Multi-agent Reinforcement LearningAttacking and Defending Deep Reinforcement Learning Policies
Recent studies have shown that deep reinforcement learning (DRL) policies are vulnerable to adversarial attacks, which raise concerns about applications of DRL to safety-critical systems. In this work, we adopt a princip…
Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Robust Android Malware Detection System against Adversarial Attacks using Q-Learning
The current state-of-the-art Android malware detection systems are based on machine learning and deep learning models. Despite having superior performance, these models are susceptible to adversarial attacks. Therefore i…
Adversarial DefenseAndroid Malware DetectionBIG-bench Machine LearningMalware Detection+4