paper-with-me

홈 › Papers

An alignment safety case sketch based on debate

2025-05-06 · Marie Davidsen Buhl, Jacob Pfau, Benjamin Hilton, Geoffrey Irving

If AI systems match or exceed human capabilities on a wide range of tasks, it may become difficult for humans to efficiently judge their actions -- making it hard to use human feedback to steer them towards desirable traits. One proposed solution is to leverage another superhuman system to point out flaws in the system's outputs via a debate. This paper outlines the value of debate for AI safety, as well as the assumptions and further research required to make debate work. It does so by sketching an ``alignment safety case'' -- an argument that an AI system will not autonomously take actions which could lead to egregious harm, despite being able to do so. The sketch focuses on the risk of an AI R\&D agent inside an AI company sabotaging research, for example by producing false results. To prevent this, the agent is trained via debate, subject to exploration guarantees, to teach the system to be honest. Honesty is maintained throughout deployment via online training. The safety case rests on four key claims: (1) the agent has become good at the debate game, (2) good performance in the debate game implies that the system is mostly honest, (3) the system will not become significantly less honest during deployment, and (4) the deployment context is tolerant of some errors. We identify open research problems that, if solved, could render this a compelling argument that an AI system is safe.

📄 PDF Abstract BibTeX arXiv:2505.03989

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases

2026-03-08 · Shaun Feakins, Ibrahim Habli, Phillip Morgan arxiv

This paper contributes to the nascent debate around safety cases for frontier AI systems. Safety cases are structured, defensible arguments that a system is acceptably safe to deploy in a given context. Historically, the…

Towards evaluations-based safety cases for AI scheming

2024-10-29 · Mikita Balesni, Marius Hobbhahn, David Lindner, Alexander Meinke 외

We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through scheming. Scheming is a potential threat m…

A sketch of an AI control safety case

2025-01-28 · Tomek Korbak, Joshua Clymer, Benjamin Hilton, Buck Shlegeris 외

As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how developers could construct a "control safety …

HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate

2025-12-09 · Shenzhe Zhu arxiv

Large language models (LLMs) are equipped with safety mechanisms to detect and block harmful queries, yet current alignment approaches primarily focus on overtly dangerous content and overlook more subtle threats. Howeve…

Scalable AI Safety via Doubly-Efficient Debate

2023-11-23 · Jonah Brown-Cohen, Geoffrey Irving, Georgios Piliouras

The emergence of pre-trained AI systems with powerful capabilities across a diverse and ever-increasing set of complex domains has raised a critical challenge for AI safety as tasks can become too complicated for humans …