paper-with-me

Papers

Parametrically Retargetable Decision-Makers Tend To Seek Power

2022-06-27 · Alexander Matt Turner, Prasad Tadepalli

If capable AI agents are generally incentivized to seek power in service of the objectives we specify for them, then these systems will pose enormous risks, in addition to enormous benefits. In fully observable environments, most reward functions have an optimal policy which seeks power by keeping options open and staying alive. However, the real world is neither fully observable, nor must trained agents be even approximately reward-optimal. We consider a range of models of AI decision-making, from optimal, to random, to choices informed by learning and interacting with an environment. We discover that many decision-making functions are retargetable, and that retargetability is sufficient to cause power-seeking tendencies. Our functional criterion is simple and broad. We show that a range of qualitatively dissimilar decision-making procedures incentivize agents to seek power. We demonstrate the flexibility of our results by reasoning about learned policy incentives in Montezuma's Revenge. These results suggest a safety risk: Eventually, retargetable training procedures may train real-world agents which seek power over humans.

📄 PDF Abstract BibTeX arXiv:2206.13477

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingMontezuma's Revenge

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Effects of AI Feedback on Learning, the Skill Gap, and Intellectual Diversity

2024-09-27 · Christoph Riedl, Eric Bogert

Can human decision-makers learn from AI feedback? Using data on 52,000 decision-makers from a large online chess platform, we investigate how their AI use affects three interrelated long-term outcomes: Learning, skill ga…

Diversity

On Avoiding Power-Seeking by Artificial Intelligence

2022-06-23 · Alexander Matt Turner

We do not know how to align a very intelligent AI agent's behavior with human interests. I investigate whether -- absent a full solution to this AI alignment problem -- we can build smart AI agents which have limited imp…

Decision Making

ABI Approach: Automatic Bias Identification in Decision-Making Under Risk based in an Ontology of Behavioral Economics

2024-05-22 · Eduardo da C. Ramos, Maria Luiza M. Campos, Fernanda Baião

Organizational decision-making is crucial for success, yet cognitive biases can significantly affect risk preferences, leading to suboptimal outcomes. Risk seeking preferences for losses, driven by biases such as loss av…

Decision MakingSystematic Literature Review

Risk Preferences in Time Lotteries

2021-08-18 · Yonatan Berman, Mark Kirstein

An important but understudied question in economics is how people choose when facing uncertainty in the timing of events. Here we study preferences over time lotteries, in which the payment amount is certain but the paym…

Calibrating Predictions to Decisions: A Novel Approach to Multi-Class Calibration

2021-07-12 · NeurIPS 2021 12 · Shengjia Zhao, Michael P. Kim, Roshni Sahoo, Tengyu Ma 외

When facing uncertainty, decision-makers want predictions they can trust. A machine learning provider can convey confidence to decision-makers by guaranteeing their predictions are distribution calibrated -- amongst the …

Decision Making