paper-with-me

홈 › Papers

On Corrigibility and Alignment in Multi Agent Games

2025-01-09 · Edmund Dable-Heath, Boyko Vodenicharski, James Bishop

Corrigibility of autonomous agents is an under explored part of system design, with previous work focusing on single agent systems. It has been suggested that uncertainty over the human preferences acts to keep the agents corrigible, even in the face of human irrationality. We present a general framework for modelling corrigibility in a multi-agent setting as a 2 player game in which the agents always have a move in which they can ask the human for supervision. This is formulated as a Bayesian game for the purpose of introducing uncertainty over the human beliefs. We further analyse two specific cases. First, a two player corrigibility game, in which we want corrigibility displayed in both agents for both common payoff (monotone) games and harmonic games. Then we investigate an adversary setting, in which one agent is considered to be a defending' agent and the other an adversary'. A general result is provided for what belief over the games and human rationality the defending agent is required to have to induce corrigibility.

📄 PDF Abstract BibTeX arXiv:2501.05360

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use

2026-05-29 · Jeremy Tien, Abishek Anand, Yu-Rou Tuan, Yuchen Shen 외 arxiv

As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safety considerations surrounding these agents become paramount. Although m…

Human Control: Definitions and Algorithms

2023-05-31 · Ryan Carey, Tom Everitt

How can humans stay in control of advanced artificial intelligence systems? One proposal is corrigibility, which requires the agent to follow the instructions of a human overseer, without inappropriately influencing them…

Corrigibility with Utility Preservation

2019-08-05 · Koen Holtman

Corrigibility is a safety property for artificially intelligent agents. A corrigible agent will not resist attempts by authorized parties to alter the goals and constraints that were encoded in the agent when it was firs…

Core Safety Values for Provably Corrigible Agents

2025-07-28 · Aran Nayebi arxiv

We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework consists of five *structurally separate* uti…

Corrigibility Transformation: Constructing Goals That Accept Updates

2025-10-17 · Rubi Hudson arxiv

For an AI's training process to successfully impart a desired goal, it is important that the AI does not attempt to resist the training. However, partially learned goals will often incentivize an AI to avoid further goal…