paper-with-me

Papers

Value alignment: a formal approach

2021-10-18 · Carles Sierra, Nardine Osman, Pablo Noriega, Jordi Sabater-Mir, Antoni Perelló

principles that should govern autonomous AI systems. It essentially states that a system's goals and behaviour should be aligned with human values. But how to ensure value alignment? In this paper we first provide a formal model to represent values through preferences and ways to compute value aggregations; i.e. preferences with respect to a group of agents and/or preferences with respect to sets of values. Value alignment is then defined, and computed, for a given norm with respect to a given value through the increase/decrease that it results in the preferences of future states of the world. We focus on norms as it is norms that govern behaviour, and as such, the alignment of a given system with a given value will be dictated by the norms the system follows.

📄 PDF Abstract BibTeX arXiv:2110.09240

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Measuring Value Alignment

2023-12-23 · Fazl Barez, Philip Torr

As artificial intelligence (AI) systems become increasingly integrated into various domains, ensuring that they align with human values becomes critical. This paper introduces a novel formalism to quantify the alignment …

Autonomous VehiclesRecommendation Systems

Concept Alignment as a Prerequisite for Value Alignment

2023-10-30 · Sunayana Rane, Mark Ho, Ilia Sucholutsky, Thomas L. Griffiths

Value alignment is essential for building AI systems that can safely and reliably interact with people. However, what a person values -- and is even capable of valuing -- depends on the concepts that they are currently u…

Concept Alignment

Value Alignment Verification

2020-10-16 · NeurIPS Workshop HAMLETS 2020 12 · Anonymous

As humans interact with autonomous agents to perform increasingly complicated, potentially risky tasks, it is important that humans can verify these agents' trustworthiness and efficiently evaluate their performance and …

Autonomous Driving

Value-Aware Multiagent Systems

2025-12-14 · Nardine Osman arxiv

This paper introduces the concept of value awareness in AI, which goes beyond the traditional value-alignment problem. Our definition of value awareness presents us with a concise and simplified roadmap for engineering v…

Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences

2025-11-03 · Joshua Ashkinaze, Hua Shen, Saipranav Avula, Eric Gilbert 외 arxiv

We introduce the Deep Value Benchmark (DVB), an evaluation framework that directly tests whether large language models (LLMs) learn fundamental human values or merely surface-level preferences. This distinction is critic…