paper-with-me

홈 › Papers

How adversarial attacks can disrupt seemingly stable accurate classifiers

2023-09-07 · Oliver J. Sutton, Qinghua Zhou, Ivan Y. Tyukin, Alexander N. Gorban, Alexander Bastounis, Desmond J. Higham

Adversarial attacks dramatically change the output of an otherwise accurate learning system using a seemingly inconsequential modification to a piece of input data. Paradoxically, empirical evidence indicates that even systems which are robust to large random perturbations of the input data remain susceptible to small, easily constructed, adversarial perturbations of their inputs. Here, we show that this may be seen as a fundamental feature of classifiers working with high dimensional input data. We introduce a simple generic and generalisable framework for which key behaviours observed in practical systems arise with high probability -- notably the simultaneous susceptibility of the (otherwise accurate) model to easily constructed adversarial attacks, and robustness to random perturbations of the input data. We confirm that the same phenomena are directly observed in practical neural networks trained on standard image classification problems, where even large additive random noise fails to trigger the adversarial instability of the network. A surprising takeaway is that even small margins separating a classifier's decision surface from training and testing data can hide adversarial susceptibility from being detected using randomly sampled perturbations. Counterintuitively, using additive noise during training or testing is therefore inefficient for eradicating or detecting adversarial examples, and more demanding adversarial training is required.

📄 PDF Abstract BibTeX arXiv:2309.03665

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage Classification

Similar Papers 제목 키워드 기반

Reinforcement Learning Disrupts Gradient-Based Adversarial Optimization

2026-06-10 · Xinhai Zou, Chang Zhao, Alireza Aghabagherloo, Dave Singelée 외 arxiv

Gradient-based adversarial attacks remain a dominant threat to deep neural networks (DNNs), as they exploit gradient information to efficiently optimize adversarial perturbations. To address this, we investigate whether …

Reinforcement Learning

Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models

2026-02-10 · Xinwei Zhang, Li Bai, Tianwei Zhang, Youqian Zhang 외 arxiv

Large vision-language models (LVLMs) have achieved impressive performance across multimodal tasks, but their reliance on visual inputs exposes them to adversarial threats. Encoder-based attacks provide an efficient alter…

Disrupting Adversarial Transferability in Deep Neural Networks

2021-08-27 · Christopher Wiedeman, Ge Wang

Adversarial attack transferability is well-recognized in deep learning. Prior work has partially explained transferability by recognizing common adversarial subspaces and correlations between decision boundaries, but lit…

Adversarial AttackFeature Correlation

Adversarial Detection with a Dynamically Stable System

2024-11-11 · Xiaowei Long, Jie Lin, Xiangyuan Yang

Adversarial detection is designed to identify and reject maliciously crafted adversarial examples(AEs) which are generated to disrupt the classification of target models. Presently, various input transformation-based met…

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

2026-06-11 · Zihao Wang, Yiming Li, Yutong Wu, Zheyu Liu 외 arxiv

Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content and execute actions with direct consequences. This makes them vulnerab…