NaturalAdversaries: Can Naturalistic Adversaries Be as Effective as Artificial Adversaries?
While a substantial body of prior work has explored adversarial example generation for natural language understanding tasks, these examples are often unrealistic and diverge from the real-world data distributions. In this work, we introduce a two-stage adversarial example generation framework (NaturalAdversaries), for designing adversaries that are effective at fooling a given classifier and demonstrate natural-looking failure cases that could plausibly occur during in-the-wild deployment of the models. At the first stage a token attribution method is used to summarize a given classifier's behaviour as a function of the key tokens in the input. In the second stage a generative model is conditioned on the key tokens from the first stage. NaturalAdversaries is adaptable to both black-box and white-box adversarial attacks based on the level of access to the model parameters. Our results indicate these adversaries generalize across domains, and offer insights for future research on improving robustness of neural text classification models.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language Understandingtext-classificationText ClassificationSimilar Papers 제목 키워드 기반
Enhancing Resilience of Deep Learning Networks by Means of Transferable Adversaries
Artificial neural networks in general and deep learning networks in particular established themselves as popular and powerful machine learning algorithms. While the often tremendous sizes of these networks are beneficial…
AI-Driven Adaptive Adversaries and the Erosion of Cryptographic Trust in Public Key Systems
This paper examines the erosion of Public Key Cryptography (PKC) security under adaptive adversarial optimisation driven by artificial intelligence. The problem addressed is the growing mismatch between algorithm-centric…
On the power of adaptivity in statistical adversaries
We study a fundamental question concerning adversarial noise models in statistical problems where the algorithm receives i.i.d. draws from a distribution $\mathcal{D}$. The definitions of these adversaries specify the ty…
AllCombining Adversaries with Anti-adversaries in Training
Adversarial training is an effective learning technique to improve the robustness of deep neural networks. In this study, the influence of adversarial training on deep learning models in terms of fairness, robustness, an…
FairnessMeta-LearningImproving Local Effectiveness for Global Robustness Training
Despite its increasing popularity, deep neural networks are easily fooled. Toalleviate this deficiency, researchers are actively developing new training strategies,which encourage models that are robust to small input…