paper-with-me

홈 › Papers

Defending Observation Attacks in Deep Reinforcement Learning via Detection and Denoising

2022-06-14 · Zikang Xiong, Joe Eappen, He Zhu, Suresh Jagannathan

Neural network policies trained using Deep Reinforcement Learning (DRL) are well-known to be susceptible to adversarial attacks. In this paper, we consider attacks manifesting as perturbations in the observation space managed by the external environment. These attacks have been shown to downgrade policy performance significantly. We focus our attention on well-trained deterministic and stochastic neural network policies in the context of continuous control benchmarks subject to four well-studied observation space adversarial attacks. To defend against these attacks, we propose a novel defense strategy using a detect-and-denoise schema. Unlike previous adversarial training approaches that sample data in adversarial scenarios, our solution does not require sampling data in an environment under attack, thereby greatly reducing risk during training. Detailed experimental results show that our technique is comparable with state-of-the-art adversarial training approaches.

📄 PDF Abstract BibTeX arXiv:2206.07188

Code (1)

ZikangXiong/rl-detect-and-denoise-defense 공식 구현 pytorch

Tasks

continuous-controlContinuous ControlDeep Reinforcement LearningDenoisingreinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Ada3Diff: Defending against 3D Adversarial Point Clouds via Adaptive Diffusion

2022-11-29 · Kui Zhang, Hang Zhou, Jie Zhang, Qidong Huang 외

Deep 3D point cloud models are sensitive to adversarial attacks, which poses threats to safety-critical applications such as autonomous driving. Robust training and defend-by-denoising are typical strategies for defendin…

Autonomous DrivingDenoising

A Computationally Efficient Method for Defending Adversarial Deep Learning Attacks

2019-06-13 · Rajeev Sahay, Rehana Mahfuz, Aly El Gamal

The reliance on deep learning algorithms has grown significantly in recent years. Yet, these models are highly vulnerable to adversarial attacks, which introduce visually imperceptible perturbations into testing data to …

Adversarial AttackDeep LearningDenoisingDimensionality Reduction

Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing

2026-01-15 · Yinzhi Zhao, Ming Wang, Shi Feng, Xiaocui Yang 외 arxiv

Large language models (LLMs) have achieved impressive performance across natural language tasks and are increasingly deployed in real-world applications. Despite extensive safety alignment efforts, recent studies show th…

RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models

2021-10-15 · EMNLP 2021 11 · Wenkai Yang, Yankai Lin, Peng Li, Jie zhou 외

Backdoor attacks, which maliciously control a well-trained model's outputs of the instances with specific triggers, are recently shown to be serious threats to the safety of reusing deep neural networks (DNNs). In this w…

Sentiment Analysis

Graph Defense Diffusion Model

2025-01-20 · Xin He, Wenqi Fan, Yili Wang, Chengyi Liu 외

Graph Neural Networks (GNNs) demonstrate significant potential in various applications but remain highly vulnerable to adversarial attacks, which can greatly degrade their performance. Existing graph purification methods…

Denoisingmodel