paper-with-me

Papers

Evaluating Frontier Models for Stealth and Situational Awareness

2025-05-02 · Mary Phuong, Roland S. Zimmermann, Ziyue Wang, David Lindner, Victoria Krakovna, Sarah Cogan, Allan Dafoe, Lewis Ho, Rohin Shah

Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavior could be very hard to detect, and if present in future advanced systems, could pose severe loss of control risk. It is therefore important for AI developers to rule out harm from scheming prior to model deployment. In this paper, we present a suite of scheming reasoning evaluations measuring two types of reasoning capabilities that we believe are prerequisites for successful scheming: First, we propose five evaluations of ability to reason about and circumvent oversight (stealth). Second, we present eleven evaluations for measuring a model's ability to instrumentally reason about itself, its environment and its deployment (situational awareness). We demonstrate how these evaluations can be used as part of a scheming inability safety case: a model that does not succeed on these evaluations is almost certainly incapable of causing severe harm via scheming in real deployment. We run our evaluations on current frontier models and find that none of them show concerning levels of either situational awareness or stealth.

📄 PDF Abstract BibTeX arXiv:2505.01420

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating whether AI models would sabotage AI safety research

2026-04-27 · Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D'Cruz 외 arxiv

We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two complementary evaluations to four Claude m…

Taken out of context: On measuring situational awareness in LLMs

2023-09-01 · Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann 외

We aim to better understand the emergence of `situational awareness' in large language models (LLMs). A model is situationally aware if it's aware that it's a model and can recognize whether it's currently in testing or …

Data AugmentationIn-Context Learning

Small-Scale Testbed for Evaluating C-V2X Applications on 5G Cellular Networks

2024-05-09 · Kaj Munhoz Arfvidsson, Kleio Fragkedaki, Frank J. Jiang, Vandana Narri 외

In this work, we present a small-scale testbed for evaluating the real-life performance of cellular V2X (C-V2X) applications on 5G cellular networks. Despite the growing interest and rapid technology development for V2X …

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

2025-08-29 · Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano 외 arxiv

Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems beh…

Instruction Following

Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs

2024-07-05 · Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan 외

AI assistants such as ChatGPT are trained to respond to users by saying, "I am a large language model". This raises questions. Do such models know that they are LLMs and reliably act on this knowledge? Are they aware of …

General KnowledgeInstruction FollowingLanguage ModellingLarge Language Model+2