paper-with-me

홈 › Papers

How Wrong Am I? - Studying Adversarial Examples and their Impact on Uncertainty in Gaussian Process Machine Learning Models

2017-11-17 · Kathrin Grosse, David Pfaff, Michael Thomas Smith, Michael Backes

Machine learning models are vulnerable to Adversarial Examples: minor perturbations to input samples intended to deliberately cause misclassification. Current defenses against adversarial examples, especially for Deep Neural Networks (DNN), are primarily derived from empirical developments, and their security guarantees are often only justified retroactively. Many defenses therefore rely on hidden assumptions that are subsequently subverted by increasingly elaborate attacks. This is not surprising: deep learning notoriously lacks a comprehensive mathematical framework to provide meaningful guarantees. In this paper, we leverage Gaussian Processes to investigate adversarial examples in the framework of Bayesian inference. Across different models and datasets, we find deviating levels of uncertainty reflect the perturbation introduced to benign samples by state-of-the-art attacks, including novel white-box attacks on Gaussian Processes. Our experiments demonstrate that even unoptimized uncertainty thresholds already reject adversarial examples in many scenarios. Comment: Thresholds can be broken in a modified attack, which was done in arXiv:1812.02606 (The limitations of model uncertainty in adversarial settings).

📄 PDF Abstract BibTeX arXiv:1711.06598

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian InferenceGaussian Processes

Similar Papers 제목 키워드 기반

Understanding Misclassifications by Attributes

2019-10-15 · Sadaf Gulshad, Zeynep Akata, Jan Hendrik Metzen, Arnold Smeulders

In this paper, we aim to understand and explain the decisions of deep neural networks by studying the behavior of predicted attributes when adversarial examples are introduced. We study the changes in attributes for clea…

Diffusion-Based Adversarial Purification for Speaker Verification

2023-10-22 · Yibo Bai, Xiao-Lei Zhang, Xuelong Li

Recently, automatic speaker verification (ASV) based on deep learning is easily contaminated by adversarial attacks, which is a new type of attack that injects imperceptible perturbations to audio signals so as to make A…

Adversarial PurificationDenoisingSpeaker Verification

Adversarial Examples Make Strong Poisons

2021-06-21 · NeurIPS 2021 12 · Liam Fowl, Micah Goldblum, Ping-Yeh Chiang, Jonas Geiping 외

The adversarial machine learning literature is largely partitioned into evasion attacks on testing data and poisoning attacks on training data. In this work, we show that adversarial examples, originally intended for att…

Data Poisoning

Context-aware Adversarial Attack on Named Entity Recognition

2023-09-16 · Shuguang Chen, Leonardo Neves, Thamar Solorio

In recent years, large pre-trained language models (PLMs) have achieved remarkable performance on many natural language processing benchmarks. Despite their success, prior studies have shown that PLMs are vulnerable to a…

Adversarial Attacknamed-entity-recognitionNamed Entity Recognition

Detecting Adversarial Perturbations with Saliency

2018-03-23 · Chiliang Zhang, Zhimou Yang, Zuochang Ye

In this paper we propose a novel method for detecting adversarial examples by training a binary classifier with both origin data and saliency data. In the case of image classification model, saliency simply explain how t…

ClassificationGeneral Classificationimage-classificationImage Classification