paper-with-me

홈 › Papers

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models

2024-12-02 · Cameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger, Prithviraj Singh Shahani, Fedor Ryzhenkov, Jacob Haimes, Felix Hofstätter, Teun van der Weij

Capability evaluations play a critical role in ensuring the safe deployment of frontier AI systems, but this role may be undermined by intentional underperformance or ``sandbagging.'' We present a novel model-agnostic method for detecting sandbagging behavior using noise injection. Our approach is founded on the observation that introducing Gaussian noise into the weights of models either prompted or fine-tuned to sandbag can considerably improve their performance. We test this technique across a range of model sizes and multiple-choice question benchmarks (MMLU, AI2, WMDP). Our results demonstrate that noise injected sandbagging models show performance improvements compared to standard models. Leveraging this effect, we develop a classifier that consistently identifies sandbagging behavior. Our unsupervised technique can be immediately implemented by frontier labs or regulatory bodies with access to weights to improve the trustworthiness of capability evaluations.

📄 PDF Abstract BibTeX arXiv:2412.01784

Code (1)

camtice/sandbagdetect 공식 구현

Tasks

MMLUMultiple-choice

Similar Papers 제목 키워드 기반

Auditing Games for Sandbagging

2025-12-08 · Jordan Taylor, Sid Black, Dillon Bowen, Thomas Read 외 arxiv

Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techniques using an auditing game. First, a re…

Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging

2026-04-29 · Jon-Paul Cacioli arxiv

A predecessor pilot (Cacioli, 2026) found that Llama-3-8B implements prompted sandbagging as positional collapse rather than answer avoidance. However, fixed option ordering in MMLU-Pro left open whether this reflected a…

Sandbagging in a Simple Survival Bandit Problem

2025-09-30 · Joel Dyer, Daniel Jarne Ornia, Nicholas Bishop, Anisoara Calinescu 외 arxiv

Evaluating the safety of frontier AI systems is an increasingly important concern, helping to measure the capabilities of such models and identify risks before deployment. However, it has been recognised that if AI agent…

AI Sandbagging: Language Models can Strategically Underperform on Evaluations

2024-06-11 · Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F. Brown 외

Trustworthy capability evaluations are crucial for ensuring the safety of AI systems, and are becoming a key component of AI regulation. However, the developers of an AI system, or the AI system itself, may have incentiv…

LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring

2025-07-31 · Chloe Li, Mary Phuong, Noah Y. Siegel arxiv

Trustworthy evaluations of dangerous capabilities are increasingly crucial for determining whether an AI system is safe to deploy. One empirically demonstrated threat is sandbagging - the strategic underperformance on ev…