paper-with-me

홈 › Papers

Forcing Generative Models to Degenerate Ones: The Power of Data Poisoning Attacks

2023-12-07 · Shuli Jiang, Swanand Ravindra Kadhe, Yi Zhou, Ling Cai, Nathalie Baracaldo

Growing applications of large language models (LLMs) trained by a third party raise serious concerns on the security vulnerability of LLMs.It has been demonstrated that malicious actors can covertly exploit these vulnerabilities in LLMs through poisoning attacks aimed at generating undesirable outputs. While poisoning attacks have received significant attention in the image domain (e.g., object detection), and classification tasks, their implications for generative models, particularly in the realm of natural language generation (NLG) tasks, remain poorly understood. To bridge this gap, we perform a comprehensive exploration of various poisoning techniques to assess their effectiveness across a range of generative tasks. Furthermore, we introduce a range of metrics designed to quantify the success and stealthiness of poisoning attacks specifically tailored to NLG tasks. Through extensive experiments on multiple NLG tasks, LLMs and datasets, we show that it is possible to successfully poison an LLM during the fine-tuning stage using as little as 1\% of the total tuning data samples. Our paper presents the first systematic approach to comprehend poisoning attacks targeting NLG tasks considering a wide range of triggers and attack settings. We hope our findings will assist the AI security community in devising appropriate defenses against such threats.

📄 PDF Abstract BibTeX arXiv:2312.04748

Code (0)

등록된 구현이 없습니다.

Tasks

Data Poisoningobject-detectionObject DetectionText Generation

Similar Papers 제목 키워드 기반

Variational Optimization for Quantum Problems using Deep Generative Networks

2024-04-28 · Lingxia Zhang, Xiaodie Lin, Peidong Wang, Kaiyan Yang 외

Optimization is one of the keystones of modern science and engineering. Its applications in quantum technology and machine learning helped nurture variational quantum algorithms and generative AI respectively. We propose…

How does Lipschitz Regularization Influence GAN Training?

2018-11-23 · ECCV 2020 8 · Yipeng Qin, Niloy Mitra, Peter Wonka

Despite the success of Lipschitz regularization in stabilizing GAN training, the exact reason of its effectiveness remains poorly understood. The direct effect of $K$-Lipschitz regularization is to restrict the $L2$-norm…

Adversarial Feature Augmentation for Unsupervised Domain Adaptation

2017-11-23 · CVPR 2018 6 · Riccardo Volpi, Pietro Morerio, Silvio Savarese, Vittorio Murino

Recent works showed that Generative Adversarial Networks (GANs) can be successfully applied in unsupervised domain adaptation, where, given a labeled source dataset and an unlabeled target dataset, the goal is to train p…

Data AugmentationDomain AdaptationUnsupervised Domain Adaptation

Signal Fidelity in Degenerate and Nondegenerate Mode Parametric Amplifier Receiving Antennas

2024-03-17 · Clayton Blosser, Adrian Bauer, Jessica E. Ruyle, K. C. Kerby-Patel 외

The gain, received power bandwidth, transient characteristics, and signal fidelity of two time-varying electrically small antennas based on parametric amplifier design are studied using practical QAM signals. Results sho…

Towards Distributed Coevolutionary GANs

2018-07-21 · Abdullah Al-Dujaili, Tom Schmiedlechner, and Erik Hemberg, Una-May O'Reilly

Generative Adversarial Networks (GANs) have become one of the dominant methods for deep generative modeling. Despite their demonstrated success on multiple vision tasks, GANs are difficult to train and much research has …