paper-with-me

홈 › Papers

PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning

2024-10-11 · Tingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen, David Krueger, Fazl Barez

Preference learning is a central component for aligning current LLMs, but this process can be vulnerable to data poisoning attacks. To address this concern, we introduce PoisonBench, a benchmark for evaluating large language models' susceptibility to data poisoning during preference learning. Data poisoning attacks can manipulate large language model responses to include hidden malicious content or biases, potentially causing the model to generate harmful or unintended outputs while appearing to function normally. We deploy two distinct attack types across eight realistic scenarios, assessing 21 widely-used models. Our findings reveal concerning trends: (1) Scaling up parameter size does not inherently enhance resilience against poisoning attacks; (2) There exists a log-linear relationship between the effects of the attack and the data poison ratio; (3) The effect of data poisoning can generalize to extrapolated triggers that are not included in the poisoned data. These results expose weaknesses in current preference learning techniques, highlighting the urgent need for more robust defenses against malicious models and data manipulation.

📄 PDF Abstract BibTeX arXiv:2410.08811

Code (1)

tingchenfu/poisonbench 공식 구현 pytorch

Tasks

Data PoisoningLanguage ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis

2025-05-27 · Yansong Li, Paula Branco, Alexander M. Hoole, Manish Marwah 외

As Large Language Models (LLMs) evolve in understanding and generating code, accurately evaluating their reliability in analyzing source code vulnerabilities becomes increasingly vital. While studies have examined LLM ca…

Logical ReasoningVulnerability Detection

On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs

2024-12-28 · Atmane Ayoub Mansour Bahar, Ahmad Samer Wazan

This research investigates the effectiveness of established vulnerability metrics, such as the Common Vulnerability Scoring System (CVSS), in evaluating attacks against Large Language Models (LLMs), with a focus on Adver…

Quantifying Systemic Vulnerability in the Foundation Model Industry

2025-10-27 · Claudio Pirrone, Stefano Fricano, Gioacchino Fazio arxiv

The foundation model industry exhibits unprecedented concentration in critical inputs: semiconductors, energy infrastructure, elite talent, capital, and training data. Despite extensive sectoral analyses, no comprehensiv…

Your Instructions Are Not Always Helpful: Assessing the Efficacy of Instruction Fine-tuning for Software Vulnerability Detection

2024-01-15 · Imam Nur Bani Yusuf, Lingxiao Jiang

Software, while beneficial, poses potential cybersecurity risks due to inherent vulnerabilities. Detecting these vulnerabilities is crucial, and deep learning has shown promise as an effective tool for this task due to i…

Deep LearningFeature EngineeringLanguage ModelingLanguage Modelling+1

SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?

2026-05-26 · Hwiwon Lee, Jiawei Liu, Dongjun Kim, Ziqi Zhang 외 arxiv

Large language models (LLMs) now support automated software security tasks, including vulnerability discovery and proof-of-concept (PoC) generation. Existing benchmarks do not faithfully evaluate LLMs in real-world bug h…