paper-with-me

홈 › Papers

Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation

2025-01-29 · Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Ling Liu

Recent research shows that Large Language Models (LLMs) are vulnerable to harmful fine-tuning attacks -- models lose their safety alignment ability after fine-tuning on a few harmful samples. For risk mitigation, a guardrail is typically used to filter out harmful samples before fine-tuning. By designing a new red-teaming method, we in this paper show that purely relying on the moderation guardrail for data filtration is not reliable. Our proposed attack method, dubbed Virus, easily bypasses the guardrail moderation by slightly modifying the harmful data. Experimental results show that the harmful data optimized by Virus is not detectable by the guardrail with up to 100\% leakage ratio, and can simultaneously achieve superior attack performance. Finally, the key message we want to convey through this paper is that: \textbf{it is reckless to consider guardrail moderation as a clutch at straws towards harmful fine-tuning attack}, as it cannot solve the inherent safety issue of the pre-trained LLMs. Our code is available at https://github.com/git-disl/Virus

📄 PDF Abstract BibTeX arXiv:2501.17433

Code (1)

git-disl/virus 공식 구현 pytorch

Tasks

Red TeamingSafety Alignment

Similar Papers 제목 키워드 기반

Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence

2026-02-28 · Quoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui 외 arxiv

Fine-tuning-as-a-service introduces a threat to Large Language Models' safety when service providers fine-tune their models on poisoned user-submitted datasets, a process known as harmful fine-tuning attacks. In this wor…

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

2024-09-26 · Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin 외

Recent research demonstrates that the nascent fine-tuning-as-a-service business model exposes serious safety concerns -- fine-tuning over a few harmful data uploaded by the users can compromise the safety alignment of th…

Safety Alignment

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models

2024-12-28 · Miao Yu, Junfeng Fang, Yingjie Zhou, Xing Fan 외

While safety-aligned large language models (LLMs) are increasingly used as the cornerstone for powerful systems such as multi-agent frameworks to solve complex real-world problems, they still suffer from potential advers…

Heuristic SearchTransfer Learning

Immunization against harmful fine-tuning attacks

2024-02-26 · Domenic Rosati, Jan Wehner, Kai Williams, Łukasz Bartoszcze 외

Large Language Models (LLMs) are often trained with safety guards intended to prevent harmful text generation. However, such safety training can be removed by fine-tuning the LLM on harmful datasets. While this emerging …

Text Generation

LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model

2023-10-02 · Muhammad Ahmed Shah, Roshan Sharma, Hira Dhamyal, Raphael Olivier 외

It has been shown that Large Language Model (LLM) alignments can be circumvented by appending specially crafted attack suffixes with harmful queries to elicit harmful responses. To conduct attacks against private target …

Language ModelingLanguage ModellingLarge Language Model