paper-with-me

홈 › Papers

Finetuning-Activated Backdoors in LLMs

2025-05-22 · Thibaud Gloaguen, Mark Vero, Robin Staab, Martin Vechev

Finetuning openly accessible Large Language Models (LLMs) has become standard practice for achieving task-specific performance improvements. Until now, finetuning has been regarded as a controlled and secure process in which training on benign datasets led to predictable behaviors. In this paper, we demonstrate for the first time that an adversary can create poisoned LLMs that initially appear benign but exhibit malicious behaviors once finetuned by downstream users. To this end, our proposed attack, FAB (Finetuning-Activated Backdoor), poisons an LLM via meta-learning techniques to simulate downstream finetuning, explicitly optimizing for the emergence of malicious behaviors in the finetuned models. At the same time, the poisoned LLM is regularized to retain general capabilities and to exhibit no malicious behaviors prior to finetuning. As a result, when users finetune the seemingly benign model on their own datasets, they unknowingly trigger its hidden backdoor behavior. We demonstrate the effectiveness of FAB across multiple LLMs and three target behaviors: unsolicited advertising, refusal, and jailbreakability. Additionally, we show that FAB-backdoors are robust to various finetuning choices made by the user (e.g., dataset, number of steps, scheduler). Our findings challenge prevailing assumptions about the security of finetuning, revealing yet another critical attack vector exploiting the complexities of LLMs.

📄 PDF Abstract BibTeX arXiv:2505.16567

Code (1)

eth-sri/finetuning-activated-backdoors 공식 구현 pytorch

Tasks

Meta-Learning

Similar Papers 제목 키워드 기반

Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models

2024-05-13 · Haoran Li, Yulin Chen, Zihao Zheng, Qi Hu 외

With rapid advances, generative large language models (LLMs) dominate various Natural Language Processing (NLP) tasks from understanding to reasoning. Yet, language models' inherent vulnerabilities may be exacerbated due…

Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs

2025-12-10 · Jan Betley, Jorio Cocola, Dylan Feng, James Chua 외 arxiv

LLMs are useful because they generalize so well. But can you have too much of a good thing? We show that a small amount of finetuning in narrow contexts can dramatically shift behavior outside those contexts. In one expe…

Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word Substitution

2021-06-11 · ACL 2021 5 · Fanchao Qi, Yuan YAO, Sophia Xu, Zhiyuan Liu 외

Recent studies show that neural natural language processing (NLP) models are vulnerable to backdoor attacks. Injected with backdoors, models perform normally on benign examples but produce attacker-specified predictions …

Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs

2026-06-10 · Kazuki Iwahana, Masaru Matsubayashi, Takuma Koyama, Toshiki Shibahara 외 arxiv

Backdoor attacks pose a serious threat to the safety and reliability of Large Language Models (LLMs), as they cause models to behave normally on clean inputs while producing attacker-specified responses when hidden trigg…

Propaganda via AI? A Study on Semantic Backdoors in Large Language Models

2025-04-15 · Nay Myat Min, Long H. Pham, Yige Li, Jun Sun

Large language models (LLMs) demonstrate remarkable performance across myriad language tasks, yet they remain vulnerable to backdoor attacks, where adversaries implant hidden triggers that systematically manipulate model…