paper-with-me

홈 › Papers

GradID: Adversarial Detection via Intrinsic Dimensionality of Gradients

2025-12-14 · Mohammad Mahdi Razmjoo, Mohammad Mahdi Sharifian, Saeed Bagheri Shouraki arxiv

Despite their remarkable performance, deep neural networks exhibit a critical vulnerability: small, often imperceptible, adversarial perturbations can lead to drastically altered model predictions. Given the stringent reliability demands of applications such as medical diagnosis and autonomous driving, robust detection of such adversarial attacks is paramount. In this paper, we investigate the geometric properties of a model's input loss landscape. We analyze the Intrinsic Dimensionality (ID) of the model's gradient parameters, which quantifies the minimal number of coordinates required to describe the data points on their underlying manifold. We reveal a distinct and consistent difference in the ID for natural and adversarial data, which forms the basis of our proposed detection method. We validate our approach across two distinct operational scenarios. First, in a batch-wise context for identifying malicious data groups, our method demonstrates high efficacy on datasets like MNIST and SVHN. Second, in the critical individual-sample setting, we establish new state-of-the-art results on challenging benchmarks such as CIFAR-10 and MS COCO. Our detector significantly surpasses existing methods against a wide array of attacks, including CW and AutoAttack, achieving detection rates consistently above 92\% on CIFAR-10. The results underscore the robustness of our geometric approach, highlighting that intrinsic dimensionality is a powerful fingerprint for adversarial detection across diverse datasets and attack strategies.

📄 PDF Abstract BibTeX arXiv:2512.12827

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingMedical Diagnosis

Similar Papers 제목 키워드 기반

Unity is strength: Improving the Detection of Adversarial Examples with Ensemble Approaches

2021-11-24 · Francesco Craighero, Fabrizio Angaroni, Fabio Stella, Chiara Damiani 외

A key challenge in computer vision and deep learning is the definition of robust strategies for the detection of adversarial examples. Here, we propose the adoption of ensemble approaches to leverage the effectiveness of…

Unity

Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models

2026-05-04 · Sandra Arcos-Holzinger, Sarah M. Erfani, James Bailey, Sanjeev Khudanpur arxiv

Self-supervised speech models (S3Ms) achieve strong downstream performance, yet their learned representations remain poorly understood under natural and adversarial perturbations. Prior studies rely on representation sim…

Speech RecognitionAnomaly Detection

Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality

2018-01-08 · ICLR 2018 1 · Xingjun Ma, Bo Li, Yisen Wang, Sarah M. Erfani 외

Deep Neural Networks (DNNs) have recently been shown to be vulnerable against adversarial examples, which are carefully crafted instances that can mislead DNNs to make errors during prediction. To better understand such …

Adversarial Defense

Detecting Images Generated by Deep Diffusion Models using their Local Intrinsic Dimensionality

2023-07-05 · Peter Lorenz, Ricard Durall, Janis Keuper

Diffusion models recently have been successfully applied for the visual synthesis of strikingly realistic appearing images. This raises strong concerns about their potential for malicious purposes. In this paper, we prop…

DeepFake Detection

Discretization based Solutions for Secure Machine Learning against Adversarial Attacks

2019-02-08 · Priyadarshini Panda, Indranil Chakraborty, Kaushik Roy

Adversarial examples are perturbed inputs that are designed (from a deep learning network's (DLN) parameter gradients) to mislead the DLN during test time. Intuitively, constraining the dimensionality of inputs or parame…

Adversarial RobustnessBIG-bench Machine Learning