paper-with-me

Papers

Human Control: Definitions and Algorithms

2023-05-31 · Ryan Carey, Tom Everitt

How can humans stay in control of advanced artificial intelligence systems? One proposal is corrigibility, which requires the agent to follow the instructions of a human overseer, without inappropriately influencing them. In this paper, we formally define a variant of corrigibility called shutdown instructability, and show that it implies appropriate shutdown behavior, retention of human autonomy, and avoidance of user harm. We also analyse the related concepts of non-obstruction and shutdown alignment, three previously proposed algorithms for human control, and one new algorithm.

📄 PDF Abstract BibTeX arXiv:2305.19861

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generating Scientific Definitions with Controllable Complexity

2022-05-01 · ACL 2022 5 · Tal August, Katharina Reinecke, Noah Smith

Unfamiliar terminology and complex language can present barriers to understanding science. Natural language processing stands to help address these issues by automatically defining unfamiliar terms. We introduce a new ta…

Reranking

Generating Scientific Definitions with Controllable Complexity

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Unfamiliar terminology and complex language can present barriers to understanding science. Natural language processing stands to help address these issues by automatically defining unfamiliar terms. We introduce a new t…

Reranking

Datalism and Data Monopolies in the Era of A.I.: A Research Agenda

2023-07-16 · Catherine E. A. Mulligan, Phil Godsiff

The increasing use of data in various parts of the economic and social systems is creating a new form of monopoly: data monopolies. We illustrate that the companies using these strategies, Datalists, are challenging the …

Towards Auditing Unsupervised Learning Algorithms and Human Processes For Fairness

2022-09-20 · Ian Davidson, S. S. Ravi

Existing work on fairness typically focuses on making known machine learning algorithms fairer. Fair variants of classification, clustering, outlier detection and other styles of algorithms exist. However, an understudie…

ClassificationClusteringFairnessOutlier Detection

The invisible power of fairness. How machine learning shapes democracy

2019-03-22 · Elena Beretta, Antonio Santangelo, Bruno Lepri, Antonio Vetrò 외

Many machine learning systems make extensive use of large amounts of data regarding human behaviors. Several researchers have found various discriminatory practices related to the use of human-related machine learning sy…

BIG-bench Machine LearningCultural Vocal Bursts Intensity PredictionFairness