paper-with-me

Papers

On Avoiding Power-Seeking by Artificial Intelligence

2022-06-23 · Alexander Matt Turner

We do not know how to align a very intelligent AI agent's behavior with human interests. I investigate whether -- absent a full solution to this AI alignment problem -- we can build smart AI agents which have limited impact on the world, and which do not autonomously seek power. In this thesis, I introduce the attainable utility preservation (AUP) method. I demonstrate that AUP produces conservative, option-preserving behavior within toy gridworlds and within complex environments based off of Conway's Game of Life. I formalize the problem of side effect avoidance, which provides a way to quantify the side effects an agent had on the world. I also give a formal definition of power-seeking in the context of AI agents and show that optimal policies tend to seek power. In particular, most reward functions have optimal policies which avoid deactivation. This is a problem if we want to deactivate or correct an intelligent agent after we have deployed it. My theorems suggest that since most agent goals conflict with ours, the agent would very probably resist correction. I extend these theorems to show that power-seeking incentives occur not just for optimal decision-makers, but under a wide range of decision-making procedures.

📄 PDF Abstract BibTeX arXiv:2206.11831

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Instrumental convergence and power-seeking

2026-06-07 · David Thorstad arxiv

Recent years have seen increasing concern that artificial intelligence may soon pose an existential risk to humanity. One leading ground for concern is that artificial agents may be power-seeking, aiming to acquire power…

Artificial Intelligence: Arguments for Catastrophic Risk

2024-01-27 · Adam Bales, William D'Alessandro, Cameron Domenico Kirk-Giannini

Recent progress in artificial intelligence (AI) has drawn attention to the technology's transformative potential, including what some see as its prospects for causing large-scale harm. We review two influential arguments…

A Review of the Evidence for Existential Risk from AI via Misaligned Power-Seeking

2023-10-27 · Rose Hadshar

Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose existential risks. This…

Asymptotically Unambitious Artificial General Intelligence

2019-05-29 · Michael K. Cohen, Badri Vellambi, Marcus Hutter

General intelligence, the ability to solve arbitrary solvable problems, is supposed by many to be artificially constructible. Narrow intelligence, the ability to solve a given particularly difficult problem, has seen imp…

Self-Driving Cars

Artificial intelligence for partial differential equations in computational mechanics: A review

2024-10-21 · Yizheng Wang, Jinshuai Bai, Zhongya Lin, Qimin Wang 외

In recent years, Artificial intelligence (AI) has become ubiquitous, empowering various fields, especially integrating artificial intelligence and traditional science (AI for Science: Artificial intelligence for science)…

Operator learning