paper-with-me

홈 › Papers

Rethinking Textual Adversarial Defense for Pre-trained Language Models

2022-07-21 · Jiayi Wang, Rongzhou Bao, Zhuosheng Zhang, Hai Zhao

Although pre-trained language models (PrLMs) have achieved significant success, recent studies demonstrate that PrLMs are vulnerable to adversarial attacks. By generating adversarial examples with slight perturbations on different levels (sentence / word / character), adversarial attacks can fool PrLMs to generate incorrect predictions, which questions the robustness of PrLMs. However, we find that most existing textual adversarial examples are unnatural, which can be easily distinguished by both human and machine. Based on a general anomaly detector, we propose a novel metric (Degree of Anomaly) as a constraint to enable current adversarial attack approaches to generate more natural and imperceptible adversarial examples. Under this new constraint, the success rate of existing attacks drastically decreases, which reveals that the robustness of PrLMs is not as fragile as they claimed. In addition, we find that four types of randomization can invalidate a large portion of textual adversarial examples. Based on anomaly detector and randomization, we design a universal defense framework, which is among the first to perform textual adversarial defense without knowing the specific attack. Empirical results show that our universal defense framework achieves comparable or even higher after-attack accuracy with other specific defenses, while preserving higher original accuracy at the same time. Our work discloses the essence of textual adversarial attacks, and indicates that (1) further works of adversarial attacks should focus more on how to overcome the detection and resist the randomization, otherwise their adversarial examples would be easily detected and invalidated; and (2) compared with the unnatural and perceptible adversarial examples, it is those undetectable adversarial examples that pose real risks for PrLMs and require more attention for future robustness-enhancing strategies.

📄 PDF Abstract BibTeX arXiv:2208.10251

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial AttackAdversarial DefenseSentence

Similar Papers 제목 키워드 기반

Textual Manifold-based Defense Against Natural Language Adversarial Examples

2022-11-05 · Dang Minh Nguyen, Luu Anh Tuan

Recent studies on adversarial images have shown that they tend to leave the underlying low-dimensional data manifold, making them significantly more challenging for current models to make correct predictions. This so-cal…

Masked Language Model Based Textual Adversarial Example Detection

2023-04-18 · Xiaomei Zhang, Zhaoxi Zhang, Qi Zhong, Xufei Zheng 외

Adversarial attacks are a serious threat to the reliable deployment of machine learning models in safety-critical applications. They can misguide current models to predict incorrectly by slightly modifying the inputs. Re…

Adversarial DefenseLanguage ModelingLanguage ModellingSST-2

Large Language Model Sentinel: LLM Agent for Adversarial Purification

2024-05-24 · Guang Lin, Toshihisa Tanaka, Qibin Zhao

Over the past two years, the use of large language models (LLMs) has advanced rapidly. While these LLMs offer considerable convenience, they also raise security concerns, as LLMs are vulnerable to adversarial attacks by …

Adversarial DefenseAdversarial PurificationAdversarial RobustnessLanguage Modeling+3

Text Adversarial Purification as Defense against Adversarial Attacks

2022-03-27 · Linyang Li, Demin Song, Xipeng Qiu

Adversarial purification is a successful defense mechanism against adversarial attacks without requiring knowledge of the form of the incoming attack. Generally, adversarial purification aims to remove the adversarial pe…

Adversarial AttackAdversarial DefenseAdversarial Purification

The Best Defense is Attack: Repairing Semantics in Textual Adversarial Examples

2023-05-06 · Heng Yang, Ke Li

Recent studies have revealed the vulnerability of pre-trained language models to adversarial attacks. Existing adversarial defense techniques attempt to reconstruct adversarial examples within feature or text spaces. How…

Adversarial AttackAdversarial Defense