paper-with-me

홈 › Papers

Intriguing Properties of Adversarial Examples

2017-11-08 · ICLR 2018 1 · Ekin D. Cubuk, Barret Zoph, Samuel S. Schoenholz, Quoc V. Le

It is becoming increasingly clear that many machine learning classifiers are vulnerable to adversarial examples. In attempting to explain the origin of adversarial examples, previous studies have typically focused on the fact that neural networks operate on high dimensional data, they overfit, or they are too linear. Here we argue that the origin of adversarial examples is primarily due to an inherent uncertainty that neural networks have about their predictions. We show that the functional form of this uncertainty is independent of architecture, dataset, and training protocol; and depends only on the statistics of the logit differences of the network, which do not change significantly during training. This leads to adversarial error having a universal scaling, as a power-law, with respect to the size of the adversarial perturbation. We show that this universality holds for a broad range of datasets (MNIST, CIFAR10, ImageNet, and random data), models (including state-of-the-art deep networks, linear models, adversarially trained networks, and networks trained on randomly shuffled labels), and attacks (FGSM, step l.l., PGD). Motivated by these results, we study the effects of reducing prediction entropy on adversarial robustness. Finally, we study the effect of network architectures on adversarial sensitivity. To do this, we use neural architecture search with reinforcement learning to find adversarially robust architectures on CIFAR10. Our resulting architecture is more robust to white \emph{and} black box attacks compared to previous attempts.

📄 PDF Abstract BibTeX arXiv:1711.02846

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial RobustnessNeural Architecture SearchReinforcement Learning

Similar Papers 제목 키워드 기반

Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs

2025-06-15 · Lu Chen, Han Yang, Hu Wang, Yuxin Cao 외

Adversarial examples have attracted significant attention over the years, yet understanding their frequency-based characteristics remains insufficient. In this paper, we investigate the intriguing properties of adversari…

Adversarial Robustnessimage-classificationImage Classification

Intriguing class-wise properties of adversarial training

2021-01-01 · Qi Tian, Kun Kuang, Fei Wu, Yisen Wang

Adversarial training is one of the most effective approaches to improve model robustness against adversarial examples. However, previous works mainly focus on the overall robustness of the model, and the in-depth analysi…

Adversarial Robustness

A Frequency Perspective of Adversarial Robustness

2021-10-26 · Shishira R Maiya, Max Ehrlich, Vatsal Agarwal, Ser-Nam Lim 외

Adversarial examples pose a unique challenge for deep learning systems. Despite recent advances in both attacks and defenses, there is still a lack of clarity and consensus in the community about the true nature and unde…

Adversarial Robustness

Rethinking Model Ensemble in Transfer-based Adversarial Attacks

2023-03-16 · Huanran Chen, Yichi Zhang, Yinpeng Dong, Xiao Yang 외

It is widely recognized that deep learning models lack robustness to adversarial examples. An intriguing property of adversarial examples is that they can transfer across different models, which enables black-box attacks…

image-classificationImage ClassificationLanguage Modellingmodel+2

The Intriguing Relation Between Counterfactual Explanations and Adversarial Examples

2020-09-11 · Timo Freiesleben

The same method that creates adversarial examples (AEs) to fool image-classifiers can be used to generate counterfactual explanations (CEs) that explain algorithmic decisions. This observation has led researchers to cons…

counterfactualRelation