paper-with-me

홈 › Papers

Observation Interference in Partially Observable Assistance Games

2024-12-23 · Scott Emmons, Caspar Oesterheld, Vincent Conitzer, Stuart Russell

We study partially observable assistance games (POAGs), a model of the human-AI value alignment problem which allows the human and the AI assistant to have partial observations. Motivated by concerns of AI deception, we study a qualitatively new phenomenon made possible by partial observability: would an AI assistant ever have an incentive to interfere with the human's observations? First, we prove that sometimes an optimal assistant must take observation-interfering actions, even when the human is playing optimally, and even when there are otherwise-equivalent actions available that do not interfere with observations. Though this result seems to contradict the classic theorem from single-agent decision making that the value of perfect information is nonnegative, we resolve this seeming contradiction by developing a notion of interference defined on entire policies. This can be viewed as an extension of the classic result that the value of perfect information is nonnegative into the cooperative multiagent setting. Second, we prove that if the human is simply making decisions based on their immediate outcomes, the assistant might need to interfere with observations as a way to query the human's preferences. We show that this incentive for interference goes away if the human is playing optimally, or if we introduce a communication channel for the human to communicate their preferences to the assistant. Third, we show that if the human acts according to the Boltzmann model of irrationality, this can create an incentive for the assistant to interfere with observations. Finally, we use an experimental model to analyze tradeoffs faced by the AI assistant in practice when considering whether or not to take observation-interfering actions.

📄 PDF Abstract BibTeX arXiv:2412.17797

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols

2024-09-12 · Charlie Griffin, Louis Thomson, Buck Shlegeris, Alessandro Abate

To evaluate the safety and usefulness of deployment protocols for untrusted AIs, AI Control uses a red-teaming exercise played between a protocol designer and an adversary. This paper introduces AI-Control Games, a forma…

Decision MakingRed Teaming

On Improving Deep Reinforcement Learning for POMDPs

2017-04-26 · Pengfei Zhu, Xin Li, Pascal Poupart, Guanghui Miao

Deep Reinforcement Learning (RL) recently emerged as one of the most competitive approaches for learning in sequential decision making problems with fully observable environments, e.g., computer Go. However, very little …

Atari GamesDecision MakingDeep Reinforcement Learningreinforcement-learning+5

On Improving Deep Reinforcement Learning for POMDPs

2018-04-17 · Pengfei Zhu, Xin Li, Pascal Poupart, Guanghui Miao

Deep Reinforcement Learning (RL) recently emerged as one of the most competitive approaches for learning in sequential decision making problems with fully observable environments, e.g., computer Go. However, very little …

Atari GamesDecision MakingDeep Reinforcement Learningreinforcement-learning+5

Minimax-Optimal Policy Regret in Partially Observable Markov Games

2026-06-01 · Raman Arora arxiv

We study sequential decision-making in partially observable environments against strategic, adaptive opponents, modeled as partially observable Markov games (POMGs). The central challenge is to learn latent dynamics from…

Partially Observable Stochastic Games with Neural Perception Mechanisms

2023-10-17 · Rui Yan, Gabriel Santos, Gethin Norman, David Parker 외

Stochastic games are a well established model for multi-agent sequential decision making under uncertainty. In practical applications, though, agents often have only partial observability of their environment. Furthermor…

Decision MakingDecision Making Under UncertaintySequential Decision Making