paper-with-me

Papers

Agent Incentives: A Causal Perspective

2021-02-02 · Tom Everitt, Ryan Carey, Eric Langlois, Pedro A Ortega, Shane Legg

We present a framework for analysing agent incentives using causal influence diagrams. We establish that a well-known criterion for value of information is complete. We propose a new graphical criterion for value of control, establishing its soundness and completeness. We also introduce two new concepts for incentive analysis: response incentives indicate which changes in the environment affect an optimal decision, while instrumental control incentives establish whether an agent can influence its utility via a variable X. For both new concepts, we provide sound and complete graphical criteria. We show by example how these results can help with evaluating the safety and fairness of an AI system.

📄 PDF Abstract BibTeX arXiv:2102.01685

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Similar Papers 제목 키워드 기반

Causal Foundations of Collective Agency

2026-04-30 · Frederik Hytting Jørgensen, Sebastian Weichwald, Lewis Hammond arxiv

A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with capabilities and goals distinct from those of any individual. More ge…

The Incentives that Shape Behaviour

2020-01-20 · Ryan Carey, Eric Langlois, Tom Everitt, Shane Legg

Which variables does an agent have an incentive to control with its decision, and which variables does it have an incentive to respond to? We formalise these incentives, and demonstrate unique graphical criteria for dete…

Fairness

Understanding Agent Incentives using Causal Influence Diagrams. Part I: Single Action Settings

2019-02-26 · Tom Everitt, Pedro A. Ortega, Elizabeth Barnes, Shane Legg

Agents are systems that optimize an objective function in an environment. Together, the goal and the environment induce secondary objectives, incentives. Modeling the agent-environment interaction using causal influence …

Reinforcement LearningReinforcement Learning (RL)

Path-Specific Objectives for Safer Agent Incentives

2022-04-21 · Sebastian Farquhar, Ryan Carey, Tom Everitt

We present a general framework for training safe agents whose naive incentives are unsafe. As an example, manipulative or deceptive behaviour can improve rewards but should be avoided. Most approaches fail here: agents m…

Learning From Strategic Agents: Accuracy, Improvement, and Causality

2020-01-01 · ICML 2020 1 · Yonadav Shavit, Benjamin Edelman, Brian Axelrod

In many predictive decision-making scenarios, such as credit scoring and academic testing, a decision-maker must construct a model that accounts for agents' incentives to ``game'' their features in order to receive bette…

Decision Making