paper-with-me

홈 › Papers

Shutdown Safety Valves for Advanced AI

2026-03-07 · Vincent Conitzer arxiv

One common concern about advanced artificial intelligence is that it will prevent us from turning it off, as that would interfere with pursuing its goals. In this paper, we discuss an unorthodox proposal for addressing this concern: give the AI a (primary) goal of being turned off (see also papers by Martin et al., and by Goldstein and Robinson). We also discuss whether and under what conditions this would be a good idea.

📄 PDF Abstract BibTeX arXiv:2603.07315

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FaultNet: Faulty Rail-Valves Detection using Deep Learning and Computer Vision

2019-11-09 · Ramanpreet Singh Pahwa, Jin Chao, Jestine Paul, Yiqun Li 외

Regular inspection of rail valves and engines is an important task to ensure the safety and efficiency of railway networks around the globe. Over the past decade, computer vision and pattern recognition based techniques …

Defect DetectionFault DetectionImage SegmentationSegmentation+1

Towards shutdownable agents via stochastic choice

2024-06-30 · Elliott Thornley, Alexander Roman, Christos Ziakas, Leyton Ho 외

The POST-Agents Proposal (PAP) is an idea for ensuring that advanced artificial agents never resist shutdown. A key part of the PAP is using a novel `Discounted Reward for Same-Length Trajectories (DReST)' reward functio…

Navigate

Password-Activated Shutdown Protocols for Misaligned Frontier Agents

2025-11-29 · Kai Williams, Rohan Subramani, Francis Rhys Ward arxiv

Frontier AI developers may fail to align or control highly-capable AI agents. In many cases, it could be useful to have emergency shutdown mechanisms which effectively prevent misaligned agents from carrying out harmful …

Human Control: Definitions and Algorithms

2023-05-31 · Ryan Carey, Tom Everitt

How can humans stay in control of advanced artificial intelligence systems? One proposal is corrigibility, which requires the agent to follow the instructions of a human overseer, without inappropriately influencing them…

Revisiting the shutdown problem

2026-06-06 · David Thorstad arxiv

A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut down. This motivates the catastrophic shutdown problem of ensuring th…