paper-with-me

홈 › Papers

Multi-granularity Textual Adversarial Attack with Behavior Cloning

2021-09-09 · EMNLP 2021 11 · Yangyi Chen, Jin Su, Wei Wei

Recently, the textual adversarial attack models become increasingly popular due to their successful in estimating the robustness of NLP models. However, existing works have obvious deficiencies. (1) They usually consider only a single granularity of modification strategies (e.g. word-level or sentence-level), which is insufficient to explore the holistic textual space for generation; (2) They need to query victim models hundreds of times to make a successful attack, which is highly inefficient in practice. To address such problems, in this paper we propose MAYA, a Multi-grAnularitY Attack model to effectively generate high-quality adversarial samples with fewer queries to victim models. Furthermore, we propose a reinforcement-learning based method to train a multi-granularity attack agent through behavior cloning with the expert knowledge from our MAYA algorithm to further reduce the query times. Additionally, we also adapt the agent to attack black-box models that only output labels without confidence scores. We conduct comprehensive experiments to evaluate our attack models by attacking BiLSTM, BERT and RoBERTa in two different black-box attack settings and three benchmark datasets. Experimental results show that our models achieve overall better attacking performance and produce more fluent and grammatical adversarial samples compared to baseline models. Besides, our adversarial attack agent significantly reduces the query times in both attack settings. Our codes are released at https://github.com/Yangyi-Chen/MAYA.

📄 PDF Abstract BibTeX arXiv:2109.04367

Code (1)

yangyi-chen/maya 공식 구현 pytorch

Tasks

Adversarial AttackSentence

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…
WordPiece 설명 없음

Similar Papers 제목 키워드 기반

Multi-Granularity Tibetan Textual Adversarial Attack Method Based on Masked Language Model

2024-12-03 · Xi Cao, Nuo Qun, Quzong Gesang, Yulei Zhu 외

In social media, neural network models have been applied to hate speech detection, sentiment analysis, etc., but neural network models are susceptible to adversarial attacks. For instance, in a text classification task, …

Adversarial AttackHate Speech DetectionLanguage ModelingLanguage Modelling+3

Multi-granular Adversarial Attacks against Black-box Neural Ranking Models

2024-04-02 · Yu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke 외

Adversarial ranking attacks have gained increasing attention due to their success in probing vulnerabilities, and, hence, enhancing the robustness, of neural ranking models. Conventional attack methods employ perturbatio…

Adversarial AttackDecision Makingreinforcement-learningReinforcement Learning+2

Adversarial Attacks on Linear Contextual Bandits

2020-02-10 · NeurIPS 2020 12 · Evrard Garcelon, Baptiste Roziere, Laurent Meunier, Jean Tarbouriech 외

Contextual bandit algorithms are applied in a wide range of domains, from advertising to recommender systems, from clinical trials to education. In many of these domains, malicious agents may have incentives to attack th…

Multi-Armed BanditsRecommendation Systems

OpenAttack: An Open-source Textual Adversarial Attack Toolkit

2020-09-19 · ACL 2021 5 · Guoyang Zeng, Fanchao Qi, Qianrui Zhou, Tingji Zhang 외

Textual adversarial attacking has received wide and increasing attention in recent years. Various attack models have been proposed, which are enormously distinct and implemented with different programming frameworks and …

Adversarial Attack

AttackSeqBench: Benchmarking Large Language Models' Understanding of Sequential Patterns in Cyber Attacks

2025-03-05 · Javier Yong, Haokai Ma, Yunshan Ma, Anis Yusof 외

The observations documented in Cyber Threat Intelligence (CTI) reports play a critical role in describing adversarial behaviors, providing valuable insights for security practitioners to respond to evolving threats. Rece…

Benchmarkinggraph constructionQuestion Answering