paper-with-me

Papers

Revisiting the shutdown problem

2026-06-06 · David Thorstad arxiv

A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut down. This motivates the catastrophic shutdown problem of ensuring that agents can be shut down before causing an existential catastrophe. A range of arguments and theorems are offered to suggest that solving the catastrophic shutdown problem is difficult, bolstering arguments for existential risk and motivating a search for solutions to the catastrophic shutdown problem. This paper argues for two conclusions. First, existing arguments require substantial additional assumptions and evidence to play the desired role in motivating existential risk concerns. Second, concern for the catastrophic shutdown problem has led to technical solutions that impose a high safety tax on model performance. These results redirect research towards the identified assumptions and provide guidance for technical alignment strategies.

📄 PDF Abstract BibTeX arXiv:2606.08296

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Shutdown Problem: An AI Engineering Puzzle for Decision Theorists

2024-03-07 · Elliott Thornley

I explain the shutdown problem: the problem of designing artificial agents that (1) shut down when a shutdown button is pressed, (2) don't try to prevent or cause the pressing of the shutdown button, and (3) otherwise pu…

Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities

2025-12-29 · Alessio Benavoli, Alessandro Facchini, Marco Zaffalon arxiv

How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designi…

Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs

2025-09-13 · Jeremy Schlatter, Benjamin Weinstein-Raun, Jeffrey Ladish arxiv

In experiments spanning more than 100,000 trials across thirteen large language models, we show that several state-of-the-art models presented with a simple task (including Grok 4, GPT-5, and Gemini 2.5 Pro) sometimes ac…

The Partially Observable Off-Switch Game

2024-11-25 · Andrew Garber, Rohan Subramani, Linus Luu, Mark Bedaywi 외

A wide variety of goals could cause an AI to disable its off switch because "you can't fetch the coffee if you're dead" (Russell 2019). Prior theoretical work on this shutdown problem assumes that humans know everything …

Evaluating Shutdown Avoidance of Language Models in Textual Scenarios

2023-07-03 · Teun van der Weij, Simon Lermen, Leon Lang

Recently, there has been an increase in interest in evaluating large language models for emergent and dangerous capabilities. Importantly, agents could reason that in some scenarios their goal is better achieved if they …