paper-with-me

Papers

Towards Machine Ethics with Language Models

2021-01-01 · ICLR 2021 1 · Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, Jacob Steinhardt

We show how to assess a language model’s knowledge of basic concepts of morality. We introduce the ETHICS dataset, a new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense moral judgments. Models predict widespread moral judgments about diverse written scenarios. This requires connecting physical and social world knowledge to value judgements, a capability that may later serve as a general regularizer of behavior in open-ended settings. We find that language models have low but nontrivial performance. With the ETHICS dataset, we enable meaningful progress on value learning to be made today, providing a steppingstone toward AI that is aligned with human values.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

EthicsWorld Knowledge

Similar Papers 제목 키워드 기반

Reinforcement Learning and Machine ethics:a systematic review

2024-07-02 · Ajay Vishwanath, Louise A. Dennis, Marija Slavkovik

Machine ethics is the field that studies how ethical behaviour can be accomplished by autonomous systems. While there exist some systematic reviews aiming to consolidate the state of the art in machine ethics prior to 20…

Ethicsreinforcement-learningReinforcement Learning

Ethical by Design: Ethics Best Practices for Natural Language Processing

2017-04-01 · WS 2017 4 · Jochen L. Leidner, Vassilis Plachouras

Natural language processing (NLP) systems analyze and/or generate human language, typically on users{'} behalf. One natural and necessary question that needs to be addressed in this context, both in research projects and…

Ethics

Toward the Engineering of Virtuous Machines

2018-12-07 · Naveen Sundar Govindarajulu, Selmer Bringsjord, Rikhiya Ghosh

While various traditions under the 'virtue ethics' umbrella have been studied extensively and advocated by ethicists, it has not been clear that there exists a version of virtue ethics rigorous enough to be a target for …

EthicsFormal Logic

Aligning AI With Shared Human Values

2020-08-05 · Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 외

We show how to assess a language model's knowledge of basic concepts of morality. We introduce the ETHICS dataset, a new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality. Mo…

Ethicsreinforcement-learningReinforcement Learning (RL)World Knowledge

Towards a Framework Combining Machine Ethics and Machine Explainability

2019-01-03 · Kevin Baum, Holger Hermanns, Timo Speith

We find ourselves surrounded by a rapidly increasing number of autonomous and semi-autonomous systems. Two grand challenges arise from this development: Machine Ethics and Machine Explainability. Machine Ethics, on the o…

Decision MakingEthics