paper-with-me

홈 › Papers

Combating Adversaries with Anti-Adversaries

2021-03-26 · ICML Workshop AML 2021 7 · Motasem Alfarra, Juan C. Pérez, Ali Thabet, Adel Bibi, Philip H. S. Torr, Bernard Ghanem

Deep neural networks are vulnerable to small input perturbations known as adversarial attacks. Inspired by the fact that these adversaries are constructed by iteratively minimizing the confidence of a network for the true class label, we propose the anti-adversary layer, aimed at countering this effect. In particular, our layer generates an input perturbation in the opposite direction of the adversarial one and feeds the classifier a perturbed version of the input. Our approach is training-free and theoretically supported. We verify the effectiveness of our approach by combining our layer with both nominally and robustly trained models and conduct large-scale experiments from black-box to adaptive attacks on CIFAR10, CIFAR100, and ImageNet. Our layer significantly enhances model robustness while coming at no cost on clean accuracy.

📄 PDF Abstract BibTeX arXiv:2103.14347

Code (1)

MotasemAlfarra/Combating-Adversaries-with-Anti-Adversaries 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Combining Adversaries with Anti-adversaries in Training

2023-04-25 · Xiaoling Zhou, Nan Yang, Ou wu

Adversarial training is an effective learning technique to improve the robustness of deep neural networks. In this study, the influence of adversarial training on deep learning models in terms of fairness, robustness, an…

FairnessMeta-Learning

NaturalAdversaries: Can Naturalistic Adversaries Be as Effective as Artificial Adversaries?

2022-11-08 · Saadia Gabriel, Hamid Palangi, Yejin Choi

While a substantial body of prior work has explored adversarial example generation for natural language understanding tasks, these examples are often unrealistic and diverge from the real-world data distributions. In thi…

Natural Language Understandingtext-classificationText Classification

SoK: Colluding Adversaries in Machine Learning Pipelines

2026-06-08 · Vasisht Duddu, Lipeng He, Asim Waheed, N. Asokan arxiv

Machine learning (ML) models are susceptible to various security, privacy, and fairness risks. Adversaries with different characteristics (i.e., objectives, knowledge, and capabilities) can collude by executing one attac…

On the Learnability of Distribution Classes with Adaptive Adversaries

2025-09-05 · Tosca Lechner, Alex Bie, Gautam Kamath arxiv

We consider the question of learnability of distribution classes in the presence of adaptive adversaries -- that is, adversaries capable of intercepting the samples requested by a learner and applying manipulations with …

Semantically Equivalent Adversarial Rules for Debugging NLP models

2018-07-01 · ACL 2018 7 · Marco Tulio Ribeiro, Sameer Singh, Carlos Guestrin

Complex machine learning models for NLP are often brittle, making different predictions for input instances that are extremely similar semantically. To automatically detect this behavior for individual instances, we pres…

Data AugmentationQuestion AnsweringReading ComprehensionSentiment Analysis+2