paper-with-me

홈 › Papers

PASA: Attack Agnostic Unsupervised Adversarial Detection using Prediction & Attribution Sensitivity Analysis

2024-04-12 · Dipkamal Bhusal, Md Tanvirul Alam, Monish K. Veerabhadran, Michael Clifford, Sara Rampazzi, Nidhi Rastogi

Deep neural networks for classification are vulnerable to adversarial attacks, where small perturbations to input samples lead to incorrect predictions. This susceptibility, combined with the black-box nature of such networks, limits their adoption in critical applications like autonomous driving. Feature-attribution-based explanation methods provide relevance of input features for model predictions on input samples, thus explaining model decisions. However, we observe that both model predictions and feature attributions for input samples are sensitive to noise. We develop a practical method for this characteristic of model prediction and feature attribution to detect adversarial samples. Our method, PASA, requires the computation of two test statistics using model prediction and feature attribution and can reliably detect adversarial samples using thresholds learned from benign samples. We validate our lightweight approach by evaluating the performance of PASA on varying strengths of FGSM, PGD, BIM, and CW attacks on multiple image and non-image datasets. On average, we outperform state-of-the-art statistical unsupervised adversarial detectors on CIFAR-10 and ImageNet by 14\% and 35\% ROC-AUC scores, respectively. Moreover, our approach demonstrates competitive performance even when an adversary is aware of the defense mechanism.

📄 PDF Abstract BibTeX arXiv:2404.10789

Code (1)

dipkamal/pasa 공식 구현 pytorch

Tasks

Autonomous DrivingSensitivity

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

EPASAD: Ellipsoid decision boundary based Process-Aware Stealthy Attack Detector

2022-04-08 · Vikas Maurya, Rachit Agarwal, Saurabh Kumar, Sandeep Kumar Shukla

Due to the importance of Critical Infrastructure (CI) in a nation's economy, they have been lucrative targets for cyber attackers. These critical infrastructures are usually Cyber-Physical Systems (CPS) such as power gri…

Anomaly DetectionIntrusion DetectionNetwork Intrusion Detection

PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks

2026-05-09 · Zhenxin Ai, Haiyun He arxiv

Watermarking for large language models (LLMs) is a promising approach for detecting LLM-generated text and enabling responsible deployment. However, existing watermarking methods are often vulnerable to semantic-invarian…

Pasadena: Perceptually Aware and Stealthy Adversarial Denoise Attack

2020-07-14 · Yupeng Cheng, Qing Guo, Felix Juefei-Xu, Wei Feng 외

Image denoising can remove natural noise that widely exists in images captured by multimedia devices due to low-quality imaging sensors, unstable image transmission processes, or low light conditions. Recent works also f…

Adversarial AttackCommon Sense ReasoningDenoisingimage-classification+2

A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP

2026-06-29 · Hodaya Krakover, Meir Yossef Levi, Eyal Gofer, Guy Gilboa arxiv

Adversarial attacks pose a challenge to the reliability of deep learning models, motivating effective detection methods. Existing techniques often rely on attack-specific assumptions, access to adversarial samples, or kn…

Adversarial Attack

Attack-Agnostic Adversarial Detection

2022-06-01 · Jiaxin Cheng, Mohamed Hussein, Jay Billa, Wael AbdAlmageed

The growing number of adversarial attacks in recent years gives attackers an advantage over defenders, as defenders must train detectors after knowing the types of attacks, and many models need to be maintained to ensure…

Adversarial AttackAdversarial Attack DetectionAnomaly Detection