paper-with-me

홈 › Papers

Attention for Adversarial Attacks: Learning from your Mistakes

2021-11-22 · AAAI Workshop AdvML 2022 2 · Florian Jaeckle, Aleksandr Agadzhanov, Jingyue Lu, M. Pawan Kumar

In order to apply Neural Networks in safety-critical settings, such as healthcare or autonomous driving, we need to be able to analyse their robustness against adversarial attacks. As complete verification is often computationally prohibitive, we rely on cheap and effective adversarial attacks to estimate their robustness. However, state-of-the-art adversarial attacks, such as the frequently used PGD attack, often require many random restarts to generate adversarial examples. Each time we perform a restart we ignore all previous unsuccessful runs. In order to alleviate this inefficiency, we propose a method that learns from its mistakes. Specifically, our method uses Graph Neural Networks (GNNs) as an attention mechanism, to greatly reduce the search space for the attacks. The architecture of the GNN is based on the neural network we are attacking, and we perform forward and backward passes though the GNN mimicking the back-propagation algorithm of PGD attacks. The GNN outputs a smaller subspace for the PGD attack to focus on. Using our method, we manage to boost the attacks' performance: the GNN increases the success rate of PGD by over 35\% on a recent published dataset used for comparing adversarial attacks, while simultaneously reducing its average computation time.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

Give Me Your Attention: Dot-Product Attention Considered Harmful for Adversarial Patch Robustness

2022-03-25 · CVPR 2022 1 · Giulio Lovisotto, Nicole Finnie, Mauricio Munoz, Chaithanya Kumar Mummadi 외

Neural architectures based on attention such as vision transformers are revolutionizing image recognition. Their main benefit is that attention allows reasoning about all parts of a scene jointly. In this paper, we show …

image-classificationImage Classificationobject-detectionObject Detection

Superclass Adversarial Attack

2022-05-29 · Soichiro Kumano, Hiroshi Kera, Toshihiko Yamasaki

Adversarial attacks have only focused on changing the predictions of the classifier, but their danger greatly depends on how the class is mistaken. For example, when an automatic driving system mistakes a Persian cat for…

Adversarial AttackMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay

2026-01-15 · Hao Wang, Yanting Wang, Hao Li, Rui Li 외 arxiv

Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to adversarial ``jailbreak'' attacks designed to bypass safety guardrails. Current safety alignment methods depend heavily on stati…

Reinforcement LearningRed Teaming

Fix your downsampling ASAP! Be natively more robust via Aliasing and Spectral Artifact free Pooling

2023-07-19 · Julia Grabinski, Janis Keuper, Margret Keuper

Convolutional neural networks encode images through a sequence of convolutions, normalizations and non-linearities as well as downsampling operations into potentially strong semantic embeddings. Yet, previous work showed…

Adversarial Attacks and Defenses: An Interpretation Perspective

2020-04-23 · Ninghao Liu, Mengnan Du, Ruocheng Guo, Huan Liu 외

Despite the recent advances in a wide spectrum of applications, machine learning models, especially deep neural networks, have been shown to be vulnerable to adversarial attacks. Attackers add carefully-crafted perturbat…

Adversarial AttackAdversarial DefenseBIG-bench Machine LearningInterpretable Machine Learning