paper-with-me

홈 › Papers

LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring

2025-07-31 · Chloe Li, Mary Phuong, Noah Y. Siegel arxiv

Trustworthy evaluations of dangerous capabilities are increasingly crucial for determining whether an AI system is safe to deploy. One empirically demonstrated threat is sandbagging - the strategic underperformance on evaluations by AI models or their developers. A promising defense is to monitor a model's chain-of-thought (CoT) reasoning, as this could reveal its intentions and plans. In this work, we measure the ability of models to sandbag on dangerous capability evaluations against a CoT monitor by prompting them to sandbag while being either monitor-oblivious or monitor-aware. We show that both frontier models and small open-sourced models can covertly sandbag against CoT monitoring 0-shot without hints. However, they cannot yet do so reliably: they bypass the monitor 16-36% of the time when monitor-aware, conditioned on sandbagging successfully. We qualitatively analyzed the uncaught CoTs to understand why the monitor failed. We reveal a rich attack surface for CoT monitoring and contribute five covert sandbagging policies generated by models. These results inform potential failure modes of CoT monitoring and may help build more diverse sandbagging model organisms.

📄 PDF Abstract BibTeX arXiv:2508.00943

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI Sandbagging: Language Models can Strategically Underperform on Evaluations

2024-06-11 · Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F. Brown 외

Trustworthy capability evaluations are crucial for ensuring the safety of AI systems, and are becoming a key component of AI regulation. However, the developers of an AI system, or the AI system itself, may have incentiv…

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models

2024-12-02 · Cameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger, Prithviraj Singh Shahani 외

Capability evaluations play a critical role in ensuring the safe deployment of frontier AI systems, but this role may be undermined by intentional underperformance or ``sandbagging.'' We present a novel model-agnostic me…

MMLUMultiple-choice

Auditing Games for Sandbagging

2025-12-08 · Jordan Taylor, Sid Black, Dillon Bowen, Thomas Read 외 arxiv

Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techniques using an auditing game. First, a re…

CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D

2025-11-13 · Francis Rhys Ward, Teun van der Weij, Hanna Gábor, Sam Martin 외 arxiv

AI systems are increasingly able to autonomously conduct realistic software engineering tasks, and may soon be deployed to automate machine learning (ML) R&D itself. Frontier AI systems may be deployed in safety-critical…

Sandbagging in a Simple Survival Bandit Problem

2025-09-30 · Joel Dyer, Daniel Jarne Ornia, Nicholas Bishop, Anisoara Calinescu 외 arxiv

Evaluating the safety of frontier AI systems is an increasingly important concern, helping to measure the capabilities of such models and identify risks before deployment. However, it has been recognised that if AI agent…