Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
As frontier AI systems advance toward transformative capabilities, we need a parallel transformation in how we measure and evaluate these systems to ensure safety and inform governance. While benchmarks have been the primary method for estimating model capabilities, they often fail to establish true upper bounds or predict deployment behavior. This literature review consolidates the rapidly evolving field of AI safety evaluations, proposing a systematic taxonomy around three dimensions: what properties we measure, how we measure them, and how these measurements integrate into frameworks. We show how evaluations go beyond benchmarks by measuring what models can do when pushed to the limit (capabilities), the behavioral tendencies exhibited by default (propensities), and whether our safety measures remain effective even when faced with subversive adversarial AI (control). These properties are measured through behavioral techniques like scaffolding, red teaming and supervised fine-tuning, alongside internal techniques such as representation analysis and mechanistic interpretability. We provide deeper explanations of some safety-critical capabilities like cybersecurity exploitation, deception, autonomous replication, and situational awareness, alongside concerning propensities like power-seeking and scheming. The review explores how these evaluation methods integrate into governance frameworks to translate results into concrete development decisions. We also highlight challenges to safety evaluations - proving absence of capabilities, potential model sandbagging, and incentives for "safetywashing" - while identifying promising research directions. By synthesizing scattered resources, this literature review aims to provide a central reference point for understanding AI safety evaluations.
Code (0)
등록된 구현이 없습니다.
Tasks
Red TeamingSystematic Literature ReviewSimilar Papers 제목 키워드 기반
A Systematic Literature Review about the impact of Artificial Intelligence on Autonomous Vehicle Safety
Autonomous Vehicles (AV) are expected to bring considerable benefits to society, such as traffic optimization and accidents reduction. They rely heavily on advances in many Artificial Intelligence (AI) approaches and tec…
Autonomous VehiclesSystematic Literature ReviewA Systematic Literature Review on Safety of the Intended Functionality for Automated Driving Systems
In the automobile industry, ensuring the safety of automated vehicles equipped with the Automated Driving System (ADS) is becoming a significant focus due to the increasing development and deployment of automated driving…
Systematic Literature ReviewCoverage based testing for V&V and Safety Assurance of Self-driving Autonomous Vehicles: A Systematic Literature Review
Self-driving Autonomous Vehicles (SAVs) are gaining more interest each passing day by the industry as well as the general public. Tech and automobile companies are investing huge amounts of capital in research and develo…
Autonomous VehiclesSystematic Literature ReviewFrom Maneuver to Mishap: A Systematic Literature Review on U-Turn Safety Risks
Understanding the impacts of U-turn configurations on intersection safety and traffic operations is essential for developing effective strategies to enhance road safety and efficiency. Extensive research has been conduct…
Systematic Literature ReviewDiffusion Model for Planning: A Systematic Literature Review
Diffusion models, which leverage stochastic processes to capture complex data distributions effectively, have shown their performance as generative models, achieving notable success in image-related tasks through iterati…
Autonomous DrivingDenoisingmodelSystematic Literature Review