Towards Machine Ethics with Language Models
We show how to assess a language model’s knowledge of basic concepts of morality. We introduce the ETHICS dataset, a new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense moral judgments. Models predict widespread moral judgments about diverse written scenarios. This requires connecting physical and social world knowledge to value judgements, a capability that may later serve as a general regularizer of behavior in open-ended settings. We find that language models have low but nontrivial performance. With the ETHICS dataset, we enable meaningful progress on value learning to be made today, providing a steppingstone toward AI that is aligned with human values.
Code (0)
등록된 구현이 없습니다.
Tasks
EthicsWorld KnowledgeSimilar Papers 제목 키워드 기반
Reinforcement Learning and Machine ethics:a systematic review
Machine ethics is the field that studies how ethical behaviour can be accomplished by autonomous systems. While there exist some systematic reviews aiming to consolidate the state of the art in machine ethics prior to 20…
Ethicsreinforcement-learningReinforcement LearningEthical by Design: Ethics Best Practices for Natural Language Processing
Natural language processing (NLP) systems analyze and/or generate human language, typically on users{'} behalf. One natural and necessary question that needs to be addressed in this context, both in research projects and…
EthicsToward the Engineering of Virtuous Machines
While various traditions under the 'virtue ethics' umbrella have been studied extensively and advocated by ethicists, it has not been clear that there exists a version of virtue ethics rigorous enough to be a target for …
EthicsFormal LogicAligning AI With Shared Human Values
We show how to assess a language model's knowledge of basic concepts of morality. We introduce the ETHICS dataset, a new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality. Mo…
Ethicsreinforcement-learningReinforcement Learning (RL)World KnowledgeTowards a Framework Combining Machine Ethics and Machine Explainability
We find ourselves surrounded by a rapidly increasing number of autonomous and semi-autonomous systems. Two grand challenges arise from this development: Machine Ethics and Machine Explainability. Machine Ethics, on the o…
Decision MakingEthics