Parametrically Retargetable Decision-Makers Tend To Seek Power
If capable AI agents are generally incentivized to seek power in service of the objectives we specify for them, then these systems will pose enormous risks, in addition to enormous benefits. In fully observable environments, most reward functions have an optimal policy which seeks power by keeping options open and staying alive. However, the real world is neither fully observable, nor must trained agents be even approximately reward-optimal. We consider a range of models of AI decision-making, from optimal, to random, to choices informed by learning and interacting with an environment. We discover that many decision-making functions are retargetable, and that retargetability is sufficient to cause power-seeking tendencies. Our functional criterion is simple and broad. We show that a range of qualitatively dissimilar decision-making procedures incentivize agents to seek power. We demonstrate the flexibility of our results by reasoning about learned policy incentives in Montezuma's Revenge. These results suggest a safety risk: Eventually, retargetable training procedures may train real-world agents which seek power over humans.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingMontezuma's RevengeMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Effects of AI Feedback on Learning, the Skill Gap, and Intellectual Diversity
Can human decision-makers learn from AI feedback? Using data on 52,000 decision-makers from a large online chess platform, we investigate how their AI use affects three interrelated long-term outcomes: Learning, skill ga…
DiversityOn Avoiding Power-Seeking by Artificial Intelligence
We do not know how to align a very intelligent AI agent's behavior with human interests. I investigate whether -- absent a full solution to this AI alignment problem -- we can build smart AI agents which have limited imp…
Decision MakingABI Approach: Automatic Bias Identification in Decision-Making Under Risk based in an Ontology of Behavioral Economics
Organizational decision-making is crucial for success, yet cognitive biases can significantly affect risk preferences, leading to suboptimal outcomes. Risk seeking preferences for losses, driven by biases such as loss av…
Decision MakingSystematic Literature ReviewRisk Preferences in Time Lotteries
An important but understudied question in economics is how people choose when facing uncertainty in the timing of events. Here we study preferences over time lotteries, in which the payment amount is certain but the paym…
Calibrating Predictions to Decisions: A Novel Approach to Multi-Class Calibration
When facing uncertainty, decision-makers want predictions they can trust. A machine learning provider can convey confidence to decision-makers by guaranteeing their predictions are distribution calibrated -- amongst the …
Decision Making