paper-with-me

Papers

Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality

2018-01-08 · ICLR 2018 1 · Xingjun Ma, Bo Li, Yisen Wang, Sarah M. Erfani, Sudanthi Wijewickrema, Grant Schoenebeck, Dawn Song, Michael E. Houle, James Bailey

Deep Neural Networks (DNNs) have recently been shown to be vulnerable against adversarial examples, which are carefully crafted instances that can mislead DNNs to make errors during prediction. To better understand such attacks, a characterization is needed of the properties of regions (the so-called 'adversarial subspaces') in which adversarial examples lie. We tackle this challenge by characterizing the dimensional properties of adversarial regions, via the use of Local Intrinsic Dimensionality (LID). LID assesses the space-filling capability of the region surrounding a reference example, based on the distance distribution of the example to its neighbors. We first provide explanations about how adversarial perturbation can affect the LID characteristic of adversarial regions, and then show empirically that LID characteristics can facilitate the distinction of adversarial examples generated using state-of-the-art attacks. As a proof-of-concept, we show that a potential application of LID is to distinguish adversarial examples, and the preliminary results show that it can outperform several state-of-the-art detection measures by large margins for five attack strategies considered in this paper across three benchmark datasets. Our analysis of the LID characteristic for adversarial regions not only motivates new directions of effective adversarial defense, but also opens up more challenges for developing new attacks to better understand the vulnerabilities of DNNs.

📄 PDF Abstract BibTeX arXiv:1801.02613

Code (1)

xingjunm/lid_adversarial_subspace_detection 공식 구현 tf

Tasks

Adversarial Defense

Similar Papers 제목 키워드 기반

On the Limitation of Local Intrinsic Dimensionality for Characterizing the Subspaces of Adversarial Examples

2018-03-26 · Pei-Hsuan Lu, Pin-Yu Chen, Chia-Mu Yu

Understanding and characterizing the subspaces of adversarial examples aid in studying the robustness of deep neural networks (DNNs) to adversarial perturbations. Very recently, Ma et al. (ICLR 2018) proposed to use loca…

On Projections to Linear Subspaces

2022-09-26 · Erik Thordsen, Erich Schubert

The merit of projecting data onto linear subspaces is well known from, e.g., dimension reduction. One key aspect of subspace projections, the maximum preservation of variance (principal component analysis), has been thor…

Dimensionality Reduction

CURVALID: Geometrically-guided Adversarial Prompt Detection

2025-03-05 · Canaan Yung, Hanxun Huang, Sarah Monazam Erfani, Christopher Leckie

Adversarial prompts capable of jailbreaking large language models (LLMs) and inducing undesirable behaviours pose a significant obstacle to their safe deployment. Current mitigation strategies rely on activating built-in…

The Space of Transferable Adversarial Examples

2017-04-11 · Florian Tramèr, Nicolas Papernot, Ian Goodfellow, Dan Boneh 외

Adversarial examples are maliciously perturbed inputs designed to mislead machine learning (ML) models at test-time. They often transfer: the same adversarial example fools more than one model. In this work, we propose…

Intrinsic Grassmann Averages for Online Linear, Robust and Nonlinear Subspace Learning

2017-02-03 · Rudrasis Chakraborty, Søren Hauberg, Baba C. Vemuri

Principal Component Analysis (PCA) and Kernel Principal Component Analysis (KPCA) are fundamental methods in machine learning for dimensionality reduction. The former is a technique for finding this approximation in fini…

Dimensionality Reduction