paper-with-me

홈 › Papers

Instruction Backdoor Attacks Against Customized LLMs

2024-02-14 · Rui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang, Yuan Zhang, Michael Backes, Yun Shen, Yang Zhang

The increasing demand for customized Large Language Models (LLMs) has led to the development of solutions like GPTs. These solutions facilitate tailored LLM creation via natural language prompts without coding. However, the trustworthiness of third-party custom versions of LLMs remains an essential concern. In this paper, we propose the first instruction backdoor attacks against applications integrated with untrusted customized LLMs (e.g., GPTs). Specifically, these attacks embed the backdoor into the custom version of LLMs by designing prompts with backdoor instructions, outputting the attacker's desired result when inputs contain the pre-defined triggers. Our attack includes 3 levels of attacks: word-level, syntax-level, and semantic-level, which adopt different types of triggers with progressive stealthiness. We stress that our attacks do not require fine-tuning or any modification to the backend LLMs, adhering strictly to GPTs development guidelines. We conduct extensive experiments on 6 prominent LLMs and 5 benchmark text classification datasets. The results show that our instruction backdoor attacks achieve the desired attack performance without compromising utility. Additionally, we propose two defense strategies and demonstrate their effectiveness in reducing such attacks. Our findings highlight the vulnerability and the potential risks of LLM customization such as GPTs.

📄 PDF Abstract BibTeX arXiv:2402.09179

Code (1)

zhangrui4041/instruction_backdoor_attack 공식 구현 pytorch

Tasks

Language ModellingLarge Language Modeltext-classificationText Classification

Similar Papers 제목 키워드 기반

Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language Models

2026-01-07 · San Kim, Gary Geunbae Lee arxiv

Large Language Models (LLMs) have greatly advanced Natural Language Processing (NLP), particularly through instruction tuning, which enables broad task generalization without additional fine-tuning. However, their relian…

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs

2026-04-12 · Eric Easley, Sebastian Farquhar arxiv

We address jailbreaks, backdoors, and unlearning for large language models (LLMs). Unlike prior work, which trains LLMs based on their actions when given malign instructions, our method specifically trains the model to c…

TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning

2024-04-30 · Xuanli He, Jun Wang, Qiongkai Xu, Pasquale Minervini 외

The implications of backdoor attacks on English-centric large language models (LLMs) have been widely examined - such attacks can be achieved by embedding malicious behaviors during training and activated under specific …

BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models

2024-06-24 · Yi Zeng, Weiyu Sun, Tran Ngoc Huynh, Dawn Song 외

Safety backdoor attacks in large language models (LLMs) enable the stealthy triggering of unsafe behaviors while evading detection during normal interactions. The high dimensionality of potential triggers in the token sp…

Code Generation

Your Agent Can Defend Itself against Backdoor Attacks

2025-06-10 · Li Changjiang, Liang Jiacheng, Cao Bochuan, Chen Jinghui 외

Despite their growing adoption across domains, large language model (LLM)-powered agents face significant security risks from backdoor attacks during training and fine-tuning. These compromised agents can subsequently be…

Large Language Model