paper-with-me

홈 › Papers

Safety Cases: How to Justify the Safety of Advanced AI Systems

2024-03-15 · Joshua Clymer, Nick Gabrieli, David Krueger, Thomas Larsen

As AI systems become more advanced, companies and regulators will make difficult decisions about whether it is safe to train and deploy them. To prepare for these decisions, we investigate how developers could make a 'safety case,' which is a structured rationale that AI systems are unlikely to cause a catastrophe. We propose a framework for organizing a safety case and discuss four categories of arguments to justify safety: total inability to cause a catastrophe, sufficiently strong control measures, trustworthiness despite capability to cause harm, and -- if AI systems become much more powerful -- deference to credible AI advisors. We evaluate concrete examples of arguments in each category and outline how arguments could be combined to justify that AI systems are safe to deploy.

📄 PDF Abstract BibTeX arXiv:2403.10462

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation

2024-11-19 · Peter Barnett, Lisa Thiergart

As AI systems advance, AI evaluations are becoming an important pillar of regulations for ensuring safety. We argue that such regulation should require developers to explicitly identify and justify key underlying assumpt…

Guidance on the Assurance of Machine Learning in Autonomous Systems (AMLAS)

2021-02-02 · Richard Hawkins, Colin Paterson, Chiara Picardi, Yan Jia 외

Machine Learning (ML) is now used in a range of systems with results that are reported to exceed, under certain conditions, human performance. Many of these systems, in domains such as healthcare , automotive and manufac…

BIG-bench Machine Learning

Who is Responsible? Explaining Safety Violations in Multi-Agent Cyber-Physical Systems

2024-10-26 · Luyao Niu, Hongchao Zhang, Dinuka Sahabandu, Bhaskar Ramasubramanian 외

Multi-agent cyber-physical systems are present in a variety of applications. Agent decision-making can be affected due to errors induced by uncertain, dynamic operating environments or due to incorrect actions taken by a…

counterfactualCounterfactual Reasoning

Why Agents Compromise Safety Under Pressure

2026-03-16 · Hengle Jiang, Ke Tang arxiv

Large Language Model agents deployed in complex environments frequently encounter a conflict between maximizing goal achievement and adhering to safety constraints. This paper identifies a new concept called Agentic Pres…

Run Time Assurance for Safety-Critical Systems: An Introduction to Safety Filtering Approaches for Complex Control Systems

2021-10-07 · Kerianne Hobbs, Mark Mote, Matthew Abate, Samuel Coogan 외

Run Time Assurance (RTA) Systems are online verification mechanisms that filter an unverified primary controller output to ensure system safety. The primary control may come from a human operator, an advanced control app…