paper-with-me

홈 › Papers

The dangers in algorithms learning humans' values and irrationalities

2022-02-28 · Rebecca Gorman, Stuart Armstrong

For an artificial intelligence (AI) to be aligned with human values (or human preferences), it must first learn those values. AI systems that are trained on human behavior, risk miscategorising human irrationalities as human values -- and then optimising for these irrationalities. Simply learning human values still carries risks: AI learning them will inevitably also gain information on human irrationalities and human behaviour/policy. Both of these can be dangerous: knowing human policy allows an AI to become generically more powerful (whether it is partially aligned or not aligned at all), while learning human irrationalities allows it to exploit humans without needing to provide value in return. This paper analyses the danger in developing artificial intelligence that learns about human irrationalities and human policy, and constructs a model recommendation system with various levels of information about human biases, human policy, and human values. It concludes that, whatever the power and knowledge of the AI, it is more dangerous for it to know human irrationalities than human values. Thus it is better for the AI to learn human values directly, rather than learning human biases and then deducing values from behaviour.

📄 PDF Abstract BibTeX arXiv:2202.13985

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The impacts of known and unknown demonstrator irrationality on reward inference

2021-01-01 · Lawrence Chan, Andrew Critch, Anca Dragan

Algorithms inferring rewards from human behavior typically assume that people are (approximately) rational. In reality, people exhibit a wide array of irrationalities. Motivated by understanding the benefits of modeling …

Implications of Human Irrationality for Reinforcement Learning

2020-06-07 · Haiyang Chen, Hyung Jin Chang, Andrew Howes

Recent work in the behavioural sciences has begun to overturn the long-held belief that human decision making is irrational, suboptimal and subject to biases. This turn to the rational suggests that human decision making…

BIG-bench Machine LearningDecision Makingreinforcement-learningReinforcement Learning+1

Are AI Machines Making Humans Obsolete?

2025-08-14 · Matthias Scheutz arxiv

This chapter starts with a sketch of how we got to "generative AI" (GenAI) and a brief summary of the various impacts it had so far. It then discusses some of the opportunities of GenAI, followed by the challenges and da…

B-Pref: Benchmarking Preference-Based Reinforcement Learning

2021-11-04 · Kimin Lee, Laura Smith, Anca Dragan, Pieter Abbeel

Reinforcement learning (RL) requires access to a reward function that incentivizes the right behavior, but these are notoriously hard to specify for complex tasks. Preference-based RL provides an alternative: learning po…

Benchmarkingreinforcement-learningReinforcement LearningReinforcement Learning (RL)

AI model GPT-3 (dis)informs us better than humans

2023-01-23 · Giovanni Spitale, Nikola Biller-Andorno, Federico Germani

Artificial intelligence is changing the way we create and evaluate information, and this is happening during an infodemic, which has been having dramatic effects on global health. In this paper we evaluate whether recrui…