paper-with-me

홈 › Papers

What AI evaluations for preventing catastrophic risks can and cannot do

2024-11-26 · Peter Barnett, Lisa Thiergart

AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can establish lower bounds on AI capabilities and assess certain misuse risks given sufficient effort from evaluators. Unfortunately, evaluations face fundamental limitations that cannot be overcome within the current paradigm. These include an inability to establish upper bounds on capabilities, reliably forecast future model capabilities, or robustly assess risks from autonomous AI systems. This means that while evaluations are valuable tools, we should not rely on them as our main way of ensuring AI systems are safe. We conclude with recommendations for incremental improvements to frontier AI safety, while acknowledging these fundamental limitations remain unsolved.

📄 PDF Abstract BibTeX arXiv:2412.08653

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Muse Spark Safety & Preparedness Report

2026-05-14 · Cristina Menghini, Peter Ney, Hamza Kwisaba, Zifan 외 arxiv

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framework, along with the evidence that informe…

Evaluating Frontier Models for Dangerous Capabilities

2024-03-20 · Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan 외

To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evaluations and pilot them on Gemini 1.0 mode…

How Catastrophic is Your LLM? Certifying Risk in Conversation

2025-10-04 · Chengxiao Wang, Isha Chaudhary, Qian Hu, Weitong Ruan 외 arxiv

Large Language Models (LLMs) can produce catastrophic responses in conversational settings that pose serious risks to public safety and security. Existing evaluations often fail to fully reveal these vulnerabilities beca…

Semantic Similarity

STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports

2025-08-13 · Tegan McCaslin, Jide Alaga, Samira Nedungadi, Seth Donoughe 외 arxiv

Evaluations of dangerous AI capabilities are important for managing catastrophic risks. Public transparency into these evaluations - including what they test, how they are conducted, and how their results inform decision…

Could AI be the Great Filter? What Astrobiology can Teach the Intelligence Community about Anthropogenic Risks

2023-05-09 · Mark M. Bailey

Where is everybody? This phrase distills the foreboding of what has come to be known as the Fermi Paradox - the disquieting idea that, if extraterrestrial life is probable in the Universe, then why have we not encountere…