paper-with-me

홈 › Papers

Universal Adversarial Triggers for Attacking and Analyzing NLP

2019-08-20 · IJCNLP 2019 11 · Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, Sameer Singh

Adversarial examples highlight model vulnerabilities and are useful for evaluation and interpretation. We define universal adversarial triggers: input-agnostic sequences of tokens that trigger a model to produce a specific prediction when concatenated to any input from a dataset. We propose a gradient-guided search over tokens which finds short trigger sequences (e.g., one word for classification and four words for language modeling) that successfully trigger the target prediction. For example, triggers cause SNLI entailment accuracy to drop from 89.94% to 0.55%, 72% of "why" questions in SQuAD to be answered "to kill american people", and the GPT-2 language model to spew racist output even when conditioned on non-racial contexts. Furthermore, although the triggers are optimized using white-box access to a specific model, they transfer to other models for all tasks we consider. Finally, since triggers are input-agnostic, they provide an analysis of global model behavior. For instance, they confirm that SNLI models exploit dataset biases and help to diagnose heuristics learned by reading comprehension models.

📄 PDF Abstract BibTeX arXiv:1908.07125

Code (1)

Eric-Wallace/universal-triggers 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingReading Comprehension

Methods 이 논문이 사용한 방법론

American 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Residual Connection 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…
Discriminative Fine-Tuning Discriminative Fine-Tuning is a fine-tuning strategy that is used for ULMFiT type models. Instead of using the same learning rate…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

Attack on Unfair ToS Clause Detection: A Case Study using Universal Adversarial Triggers

2022-11-28 · Shanshan Xu, Irina Broda, Rashid Haddad, Marco Negrini 외

Recent work has demonstrated that natural language processing techniques can support consumer protection by automatically detecting unfair clauses in the Terms of Service (ToS) Agreement. This work demonstrates that tran…

Why do universal adversarial attacks work on large language models?: Geometry might be the answer

2023-09-01 · Varshini Subhash, Anna Bialas, Weiwei Pan, Finale Doshi-Velez

Transformer based large language models with emergent capabilities are becoming increasingly ubiquitous in society. However, the task of understanding and interpreting their internal workings, in the context of adversari…

Dimensionality Reduction

MINIMAL: Mining Models for Data Free Universal Adversarial Triggers

2021-09-25 · Swapnil Parekh, Yaman Singla Kumar, Somesh Singh, Changyou Chen 외

It is well known that natural language models are vulnerable to adversarial attacks, which are mostly input-specific in nature. Recently, it has been shown that there also exist input-agnostic attacks in NLP models, call…

Natural Language Inference

Universal Adversarial Triggers Are Not Universal

2024-04-24 · Nicholas Meade, Arkil Patel, Siva Reddy

Recent work has developed optimization procedures to find token sequences, called adversarial triggers, which can elicit unsafe responses from aligned language models. These triggers are believed to be universally transf…

Exploring the Universal Vulnerability of Prompt-based Learning Paradigm

2022-04-11 · Findings (NAACL) 2022 7 · Lei Xu, Yangyi Chen, Ganqu Cui, Hongcheng Gao 외

Prompt-based learning paradigm bridges the gap between pre-training and fine-tuning, and works effectively under the few-shot setting. However, we find that this learning paradigm inherits the vulnerability from the pre-…