BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation Models
Pre-trained Natural Language Processing (NLP) models can be easily adapted to a variety of downstream language tasks. This significantly accelerates the development of language models. However, NLP models have been shown to be vulnerable to backdoor attacks, where a pre-defined trigger word in the input text causes model misprediction. Previous NLP backdoor attacks mainly focus on some specific tasks. This makes those attacks less general and applicable to other kinds of NLP models and tasks. In this work, we propose \Name, the first task-agnostic backdoor attack against the pre-trained NLP models. The key feature of our attack is that the adversary does not need prior information about the downstream tasks when implanting the backdoor to the pre-trained model. When this malicious model is released, any downstream models transferred from it will also inherit the backdoor, even after the extensive transfer learning process. We further design a simple yet effective strategy to bypass a state-of-the-art defense. Experimental results indicate that our approach can compromise a wide range of downstream NLP tasks in an effective and stealthy way.
Code (0)
등록된 구현이 없습니다.
Tasks
Backdoor AttackTransfer LearningSimilar Papers 제목 키워드 기반
Multi-target Backdoor Attacks for Code Pre-trained Models
Backdoor attacks for neural code models have gained considerable attention due to the advancement of code intelligence. However, most existing works insert triggers into task-specific data for code-related downstream tas…
Code GenerationRepresentation LearningGhostEncoder: Stealthy Backdoor Attacks with Dynamic Triggers to Pre-trained Encoders in Self-supervised Learning
Within the realm of computer vision, self-supervised learning (SSL) pertains to training pre-trained image encoders utilizing a substantial quantity of unlabeled images. Pre-trained image encoders can serve as feature ex…
Backdoor AttackImage SteganographySelf-Supervised LearningModel Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding Indistinguishability
Pre-trained models (PTMs) are widely adopted across various downstream tasks in the machine learning supply chain. Adopting untrustworthy PTMs introduces significant security risks, where adversaries can poison the model…
Backdoor AttackTask-Agnostic Detector for Insertion-Based Backdoor Attacks
Textual backdoor attacks pose significant security threats. Current detection approaches, typically relying on intermediate feature representation or reconstructing potential triggers, are task-specific and less effectiv…
named-entity-recognitionNamed Entity RecognitionQuestion AnsweringSentence+1SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
Although pre-training achieves remarkable performance, it suffers from task-agnostic backdoor attacks due to vulnerabilities in data and training mechanisms. These attacks can transfer backdoors to various downstream tas…
Backdoor AttackContrastive LearningNatural Language Understanding