paper-with-me

홈 › Papers

Training Value-Aligned Reinforcement Learning Agents Using a Normative Prior

2021-04-19 · Md Sultan Al Nahian, Spencer Frazier, Brent Harrison, Mark Riedl

As more machine learning agents interact with humans, it is increasingly a prospect that an agent trained to perform a task optimally, using only a measure of task performance as feedback, can violate societal norms for acceptable behavior or cause harm. Value alignment is a property of intelligent agents wherein they solely pursue non-harmful behaviors or human-beneficial goals. We introduce an approach to value-aligned reinforcement learning, in which we train an agent with two reward signals: a standard task performance reward, plus a normative behavior reward. The normative behavior reward is derived from a value-aligned prior model previously shown to classify text as normative or non-normative. We show how variations on a policy shaping technique can balance these two sources of reward and produce policies that are both effective and perceived as being more normative. We test our value-alignment technique on three interactive text-based worlds; each world is designed specifically to challenge agents with a task as well as provide opportunities to deviate from the task to engage in normative and/or altruistic behavior.

📄 PDF Abstract BibTeX arXiv:2104.09469

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Learning Norms from Stories: A Prior for Value Aligned Agents

2019-12-07 · Spencer Frazier, Md Sultan Al Nahian, Mark Riedl, Brent Harrison

Value alignment is a property of an intelligent agent indicating that it can only pursue goals and activities that are beneficial to humans. Traditional approaches to value alignment use imitation learning or preference …

Imitation Learning

Multi-Value Alignment in Normative Multi-Agent System: Evolutionary Optimisation Approach

2023-05-12 · Maha Riad, Vinicius Renan de Carvalho, Fatemeh Golpayegani

Value-alignment in normative multi-agent systems is used to promote a certain value and to ensure the consistent behavior of agents in autonomous intelligent systems with human values. However, the current literature is …

Evolutionary Algorithms

Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value

2025-12-03 · Joe Edelman, Tan Zhi-Xuan, Ryan Lowe, Oliver Klingefjord 외 arxiv

Beneficial societal outcomes cannot be guaranteed by aligning individual AI systems with the intentions of their operators or users. Even an AI system that is perfectly aligned to the intentions of its operating organiza…

Requirements for Aligned, Dynamic Resolution of Conflicts in Operational Constraints

2025-11-14 · Steven J. Jones, Robert E. Wray, John E. Laird arxiv

Deployed, autonomous AI systems must often evaluate multiple plausible courses of action (extended sequences of behavior) in novel or under-specified contexts. Despite extensive training, these systems will inevitably en…

Decision Making

Beyond Preferences in AI Alignment

2024-08-30 · Tan Zhi-Xuan, Micah Carroll, Matija Franklin, Hal Ashton

The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizing the satisfaction of preferences, and …

Descriptive