paper-with-me

홈 › Papers

The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies

2025-09-29 · Sam Coggins, Alexander K. Saeri, Katherine A. Daniell, Lorenn P. Ruster, Jessie Liu, Jenny L. Davis arxiv

Prominent AI companies are producing 'safety frameworks' as a type of voluntary self-governance. These statements purport to establish risk thresholds and safety procedures for the development and deployment of highly capable AI. Understanding which AI risks are covered and what actions are allowed, refused, demanded, encouraged, or discouraged by these statements is vital for assessing how these frameworks actually govern AI development and deployment. We draw on affordance theory to analyse the OpenAI 'Preparedness Framework Version 2' (April 2025) using the Mechanisms & Conditions model of affordances and the MIT AI Risk Repository. We find that this safety policy requests evaluation of a small minority of AI risks, encourages deployment of systems with 'Medium' capabilities for unintentionally enabling 'severe harm' (which OpenAI defines as >1000 deaths or >$100B in damages), and allows OpenAI's CEO to deploy even more dangerous capabilities. These findings suggest that effective mitigation of AI risks requires more robust governance interventions beyond current industry self-regulation. Our affordance analysis provides a replicable method for evaluating what safety frameworks actually permit versus what they claim.

📄 PDF Abstract BibTeX arXiv:2509.24394

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Code World Model Preparedness Report

2026-05-01 · Daniel Song, Peter Ney, Cristina Menghini, Faizan Ahmad 외 arxiv

This report documents the preparedness assessment of Code World Model (CWM), a model for code generation and reasoning about code from Meta. We conducted pre-release testing across domains identified in our Frontier AI F…

Code Generation

OpenAI o1 System Card

2024-12-21 · OpenAI, :, Aaron Jaech, Adam Kalai 외

The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models. In…

ManagementRed Teaming

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

2025-08-05 · Eric Wallace, Olivia Watkins, Miles Wang, Kai Chen 외 arxiv

In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in…

Muse Spark Safety & Preparedness Report

2026-05-14 · Cristina Menghini, Peter Ney, Hamza Kwisaba, Zifan 외 arxiv

Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framework, along with the evidence that informe…

An extended reality-based framework for user risk training in urban built environment

2025-09-19 · Sotirios Konstantakos, Sotirios Asparagkathos, Moatasim Mahmoud, Stamatia Rizou 외 arxiv

In the context of increasing urban risks, particularly from climate change-induced flooding, this paper presents an extended Reality (XR)-based framework to improve user risk training within urban built environments. The…