Turning Generative Models Degenerate: The Power of Data Poisoning Attacks
The increasing use of large language models (LLMs) trained by third parties raises significant security concerns. In particular, malicious actors can introduce backdoors through poisoning attacks to generate undesirable outputs. While such attacks have been extensively studied in image domains and classification tasks, they remain underexplored for natural language generation (NLG) tasks. To address this gap, we conduct an investigation of various poisoning techniques targeting the LLM's fine-tuning phase via prefix-tuning, a Parameter Efficient Fine-Tuning (PEFT) method. We assess their effectiveness across two generative tasks: text summarization and text completion; and we also introduce new metrics to quantify the success and stealthiness of such NLG poisoning attacks. Through our experiments, we find that the prefix-tuning hyperparameters and trigger designs are the most crucial factors to influence attack success and stealthiness. Moreover, we demonstrate that existing popular defenses are ineffective against our poisoning attacks. Our study presents the first systematic approach to understanding poisoning attacks targeting NLG tasks during fine-tuning via PEFT across a wide range of triggers and attack settings. We hope our findings will aid the AI security community in developing effective defenses against such threats.
Code (0)
등록된 구현이 없습니다.
Tasks
Data Poisoningparameter-efficient fine-tuningText GenerationText SummarizationSimilar Papers 제목 키워드 기반
Forcing Generative Models to Degenerate Ones: The Power of Data Poisoning Attacks
Growing applications of large language models (LLMs) trained by a third party raise serious concerns on the security vulnerability of LLMs.It has been demonstrated that malicious actors can covertly exploit these vulnera…
Data Poisoningobject-detectionObject DetectionText GenerationGenerating Fake Cyber Threat Intelligence Using Transformer-Based Models
Cyber-defense systems are being developed to automatically ingest Cyber Threat Intelligence (CTI) that contains semi-structured data and/or text to populate knowledge graphs. A potential risk is that fake CTI can be gene…
Data PoisoningKnowledge GraphsLanguage ModellingSentenceTurning Federated Learning Systems Into Covert Channels
Federated learning (FL) goes beyond traditional, centralized machine learning by distributing model training among a large collection of edge clients. These clients cooperatively train a global, e.g., cloud-hosted, model…
Federated LearningModel PoisoningAssociative Poisoning to Generative Machine Learning
The widespread adoption of generative models such as Stable Diffusion and ChatGPT has made them increasingly attractive targets for malicious exploitation, particularly through data poisoning. Existing poisoning attacks …
Signal Fidelity in Degenerate and Nondegenerate Mode Parametric Amplifier Receiving Antennas
The gain, received power bandwidth, transient characteristics, and signal fidelity of two time-varying electrically small antennas based on parametric amplifier design are studied using practical QAM signals. Results sho…