paper-with-me

Papers

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

2025-02-03 · Zora Che, Stephen Casper, Robert Kirk, Anirudh Satheesh, Stewart Slocum, Lev E McKinney, Rohit Gandikota, Aidan Ewart, Domenic Rosati, Zichu Wu, Zikui Cai, Bilal Chughtai, Yarin Gal, Furong Huang, Dylan Hadfield-Menell

Evaluations of large language model (LLM) risks and capabilities are increasingly being incorporated into AI risk management and governance frameworks. Currently, most risk evaluations are conducted by designing inputs that elicit harmful behaviors from the system. However, this approach suffers from two limitations. First, input-output evaluations cannot evaluate realistic risks from open-weight models. Second, the behaviors identified during any particular input-output evaluation can only lower-bound the model's worst-possible-case input-output behavior. As a complementary method for eliciting harmful behaviors, we propose evaluating LLMs with model tampering attacks which allow for modifications to latent activations or weights. We pit state-of-the-art techniques for removing harmful LLM capabilities against a suite of 5 input-space and 6 model tampering attacks. In addition to benchmarking these methods against each other, we show that (1) model resilience to capability elicitation attacks lies on a low-dimensional robustness subspace; (2) the attack success rate of model tampering attacks can empirically predict and offer conservative estimates for the success of held-out input-space attacks; and (3) state-of-the-art unlearning methods can easily be undone within 16 steps of fine-tuning. Together these results highlight the difficulty of suppressing harmful LLM capabilities and show that model tampering attacks enable substantially more rigorous evaluations than input-space attacks alone.

📄 PDF Abstract BibTeX arXiv:2502.05209

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingLarge Language Model

Similar Papers 제목 키워드 기반

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

2026-02-06 · Saad Hossain, Tom Tseng, Punya Syon Pandey, Samanvay Vajpayee 외 arxiv

As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks. Howeve…

Learning under $p$-Tampering Attacks

2017-11-10 · Saeed Mahloujifar, Dimitrios I. Diochnos, Mohammad Mahmoody

Recently, Mahloujifar and Mahmoody (TCC'17) studied attacks against learning algorithms using a special case of Valiant's malicious noise, called $p$-tampering, in which the adversary gets to change any training example …

PAC learning

Transferable Adversarial Attack on Image Tampering Localization

2023-09-19 · Yuqi Wang, Gang Cao, Zijie Lou, Haochen Zhu

It is significant to evaluate the security of existing digital image tampering localization algorithms in real-world applications. In this paper, we propose an adversarial attack scheme to reveal the reliability of such …

Adversarial Attack

Detection of Physiological Data Tampering Attacks with Quantum Machine Learning

2025-02-09 · Md. Saif Hassan Onim, Himanshu Thapliyal

The widespread use of cloud-based medical devices and wearable sensors has made physiological data susceptible to tampering. These attacks can compromise the reliability of healthcare systems which can be critical and li…

Data PoisoningQuantum Machine Learning

Resilient Linear Classification: An Approach to Deal with Attacks on Training Data

2017-08-10 · Sangdon Park, James Weimer, Insup Lee

Data-driven techniques are used in cyber-physical systems (CPS) for controlling autonomous vehicles, handling demand responses for energy management, and modeling human physiology for medical devices. These data-driven t…

Autonomous VehiclesClassificationenergy managementGeneral Classification+1