paper-with-me

홈 › Papers

Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Despite the success of recent deep learning techniques, they still perform poorly on adversarial examples with small perturbations. While gradient-based adversarial attack methods are well-explored in the field of computer vision, it is impractical to directly apply them in natural language processing due to the discrete nature of the text. To address the problem, we propose a unified framework to extend the existing gradient-based method to craft textual adversarial samples. In this framework, gradient-based continuous perturbations are added to the embedding layer and amplified in the forward propagation process. Then the final perturbed latent representations are decoded with a mask language model head to obtain potential adversarial samples. In this paper, we instantiate our framework with an attack algorithm named Textual Projected GradientDescent (T-PGD). We conduct comprehensive experiments to evaluate our framework by performing transfer black-box attacks on BERT, RoBERTa, and ALBERT on three benchmark datasets. Experimental results demonstrate that our method achieves an overall better performance and produces more fluent and grammatical adversarial samples compared to strong baseline methods. All the code and data will be made public.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial AttackLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
LAMB LAMB is a a layerwise adaptive large batch optimization technique. It provides a strategy for adapting the learning rate in large batch settings. LAMB uses…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
RoBERTa 설명 없음
Adam 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…

Similar Papers 제목 키워드 기반

Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework

2021-10-28 · Lifan Yuan, Yichi Zhang, Yangyi Chen, Wei Wei

Despite recent success on various tasks, deep learning techniques still perform poorly on adversarial examples with small perturbations. While optimization-based methods for adversarial attacks are well-explored in the f…

Adversarial AttackLanguage Modelling

TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization

2022-12-19 · Bairu Hou, Jinghan Jia, Yihua Zhang, Guanhua Zhang 외

Robustness evaluation against adversarial examples has become increasingly important to unveil the trustworthiness of the prevailing deep models in natural language processing (NLP). However, in contrast to the computer …

Adversarial DefenseAdversarial RobustnessLanguage Modelling

Bridging the Performance Gap between FGSM and PGD Adversarial Training

2020-11-07 · Tianjin Huang, Vlado Menkovski, Yulong Pei, Mykola Pechenizkiy

Deep learning achieves state-of-the-art performance in many tasks but exposes to the underlying vulnerability against adversarial examples. Across existing defense techniques, adversarial training with the projected grad…

Adversarial AttackAdversarial Robustness

NonTextual Target Attack

2025-10-03 · Xinzhe Huang, Wenjing Hu, Tianhang Zheng, Kedong Xiu 외 arxiv

Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses. However, restricting the objective as inducing f…

Bridging Adversarial Robustness and Gradient Interpretability

2019-03-27 · Beomsu Kim, Junghoon Seo, Taegyun Jeon

Adversarial training is a training scheme designed to counter adversarial attacks by augmenting the training dataset with adversarial examples. Surprisingly, several studies have observed that loss gradients from adversa…

Adversarial Robustness