paper-with-me

Papers

Metrics and methods for robustness evaluation of neural networks with generative models

2020-03-04 · Igor Buzhinsky, Arseny Nerinovsky, Stavros Tripakis

Recent studies have shown that modern deep neural network classifiers are easy to fool, assuming that an adversary is able to slightly modify their inputs. Many papers have proposed adversarial attacks, defenses and methods to measure robustness to such adversarial perturbations. However, most commonly considered adversarial examples are based on $\ell_p$-bounded perturbations in the input space of the neural network, which are unlikely to arise naturally. Recently, especially in computer vision, researchers discovered "natural" or "semantic" perturbations, such as rotations, changes of brightness, or more high-level changes, but these perturbations have not yet been systematically utilized to measure the performance of classifiers. In this paper, we propose several metrics to measure robustness of classifiers to natural adversarial examples, and methods to evaluate them. These metrics, called latent space performance metrics, are based on the ability of generative models to capture probability distributions, and are defined in their latent spaces. On three image classification case studies, we evaluate the proposed metrics for several classifiers, including ones trained in conventional and robust ways. We find that the latent counterparts of adversarial robustness are associated with the accuracy of the classifier rather than its conventional adversarial robustness, but the latter is still reflected on the properties of found latent perturbations. In addition, our novel method of finding latent adversarial perturbations demonstrates that these perturbations are often perceptually small.

📄 PDF Abstract BibTeX arXiv:2003.01993

Code (1)

igor-buzhinsky/latent-space-nn-evaluation 공식 구현 pytorch

Tasks

Adversarial Robustnessimage-classificationImage Classification

Similar Papers 제목 키워드 기반

EvalGIM: A Library for Evaluating Generative Image Models

2024-12-13 · Melissa Hall, Oscar Mañas, Reyhane Askari-Hemmat, Mark Ibrahim 외

As the use of text-to-image generative models increases, so does the adoption of automatic benchmarking methods used in their evaluation. However, while metrics and datasets abound, there are few unified benchmarking lib…

BenchmarkingDiversity

Rethinking Robustness: A New Approach to Evaluating Feature Attribution Methods

2025-12-07 · Panagiota Kiourti, Anu Singh, Preeti Duraipandian, Weichao Zhou 외 arxiv

This paper studies the robustness of feature attribution methods for deep neural networks. It challenges the current notion of attributional robustness that largely ignores the difference in the model's outputs and intro…

Enhanced Generative Model Evaluation with Clipped Density and Coverage

2025-07-02 · Nicolas Salvy, Hugues Talbot, Bertrand Thirion arxiv

Although generative models have made remarkable progress in recent years, their use in critical applications has been hindered by an inability to reliably evaluate the quality of their generated samples. Quality refers t…

MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension Metrics

2020-10-07 · EMNLP 2020 11 · Anthony Chen, Gabriel Stanovsky, Sameer Singh, Matt Gardner

Posing reading comprehension as a generation problem provides a great deal of flexibility, allowing for open-ended questions with few restrictions on possible answers. However, progress is impeded by existing generation …

Question AnsweringReading Comprehension

URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement

2024-06-07 · Wangyou Zhang, Robin Scheibler, Kohei Saijo, Samuele Cornell 외

The last decade has witnessed significant advancements in deep learning-based speech enhancement (SE). However, most existing SE research has limitations on the coverage of SE sub-tasks, data diversity and amount, and ev…

Bandwidth ExtensionDenoisingDiversitySpeech Enhancement