paper-with-me

홈 › Papers

"That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial Attacks

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Adversarial attacks are a major challenge faced by current machine learning research. These purposely crafted inputs fool even the most advanced models, precluding their deployment in safety-critical applications. Extensive research in computer vision has been carried to develop reliable defense strategies. However, the same issue remains less explored in natural language processing. Our work presents a model-agnostic detector of adversarial text examples. The approach identifies patterns in the logits of the target classifier when perturbing the input text. The proposed detector improves the current state-of-the-art performance in recognizing adversarial inputs and exhibits strong generalization capabilities across different NLP models, datasets, and word-level attacks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Text

Similar Papers 제목 키워드 기반

"That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial Attacks

2022-04-10 · Edoardo Mosca, Shreyash Agarwal, Javier Rando, Georg Groh

Adversarial attacks are a major challenge faced by current machine learning research. These purposely crafted inputs fool even the most advanced models, precluding their deployment in safety-critical applications. Extens…

Adversarial Text

“That Is a Suspicious Reaction!”: Interpreting Logits Variation to Detect NLP Adversarial Attacks

2022-05-01 · ACL 2022 5 · Edoardo Mosca, Shreyash Agarwal, Javier Rando Ramírez, Georg Groh

Adversarial attacks are a major challenge faced by current machine learning research. These purposely crafted inputs fool even the most advanced models, precluding their deployment in safety-critical applications. Extens…

Adversarial Text

Interpreting social cues to generate credible affective reactions of virtual job interviewers

2014-02-20 · Hazael Jones, Nicolas Sabouret, Ionut Damian, Tobias Baur 외

In this paper we describe a mechanism of generating credible affective reactions in a virtual recruiter during an interaction with a user. This is done using communicative performance computation based on the behaviours …

A Case for the Score: Identifying Image Anomalies using Variational Autoencoder Gradients

2019-11-28 · David Zimmerer, Jens Petersen, Simon A. A. Kohl, Klaus H. Maier-Hein

Through training on unlabeled data, anomaly detection has the potential to impact computer-aided diagnosis by outlining suspicious regions. Previous work on deep-learning-based anomaly detection has primarily focused on …

Anomaly Detection

Misinfo Reaction Frames: Reasoning about Readers’ Reactions to News Headlines

2022-05-01 · ACL 2022 5 · Saadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen 외

Even to a simple and short news headline, readers react in a multitude of ways: cognitively (e.g. inferring the writer’s intent), emotionally (e.g. feeling distrust), and behaviorally (e.g. sharing the news with their fr…

Misinformation