paper-with-me

Papers

Evaluating Frontier Models for Dangerous Capabilities

2024-03-20 · Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca Dragan, Rohin Shah, Allan Dafoe, Toby Shevlane

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 models. Our evaluations cover four areas: (1) persuasion and deception; (2) cyber-security; (3) self-proliferation; and (4) self-reasoning. We do not find evidence of strong dangerous capabilities in the models we evaluated, but we flag early warning signs. Our goal is to help advance a rigorous science of dangerous capability evaluation, in preparation for future models.

📄 PDF Abstract BibTeX arXiv:2403.13793

Code (1)

google-deepmind/dangerous-capability-evaluations

Similar Papers 제목 키워드 기반

PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach

2025-11-24 · Udari Madhushani Sehwag, Shayan Shabihi, Alex McAvoy, Vikash Sehwag 외 arxiv

Recent advances in Large Language Models (LLMs) have sparked concerns over their potential to acquire and misuse dangerous or high-risk capabilities, posing frontier risks. Current safety evaluations primarily test for w…

Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations

2024-12-19 · Paolo Bova, Alessandro Di Stefano, The Anh Han

We present a quantitative model for tracking dangerous AI capabilities over time. Our goal is to help the policy and research community visualise how dangerous capability testing can give us an early warning about approa…

Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs

2026-01-20 · Jackson Kaunismaa, Avery Griffin, John Hughes, Christina Q. Knight 외 arxiv

Model developers implement safeguards in frontier models to prevent misuse, for example, by employing classifiers to filter dangerous outputs. In this work, we demonstrate that even robustly safeguarded models can be use…

Frontier AI Regulation: Managing Emerging Risks to Public Safety

2023-07-06 · Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung 외

Advanced AI models hold the promise of tremendous benefits for humanity, but society needs to proactively manage the accompanying risks. In this paper, we focus on what we term "frontier AI" models: highly capable founda…

Sandbagging in a Simple Survival Bandit Problem

2025-09-30 · Joel Dyer, Daniel Jarne Ornia, Nicholas Bishop, Anisoara Calinescu 외 arxiv

Evaluating the safety of frontier AI systems is an increasingly important concern, helping to measure the capabilities of such models and identify risks before deployment. However, it has been recognised that if AI agent…