paper-with-me

홈 › Papers

Goal Alignment: A Human-Aware Account of Value Alignment Problem

2023-02-02 · Malek Mechergui, Sarath Sreedharan

Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objective of its users. The problem has been widely argued to be one of the central safety problems in AI. Unfortunately, most existing works in value alignment tend to focus on issues that are primarily related to the fact that reward functions are an unintuitive mechanism to specify objectives. However, the complexity of the objective specification mechanism is just one of many reasons why the user may have misspecified their objective. A foundational cause for misalignment that is being overlooked by these works is the inherent asymmetry in human expectations about the agent's behavior and the behavior generated by the agent for the specified objective. To address this lacuna, we propose a novel formulation for the value alignment problem, named goal alignment that focuses on a few central challenges related to value alignment. In doing so, we bridge the currently disparate research areas of value alignment and human-aware planning. Additionally, we propose a first-of-its-kind interactive algorithm that is capable of using information generated under incorrect beliefs about the agent, to determine the true underlying goal of the user.

📄 PDF Abstract BibTeX arXiv:2302.00813

Code (0)

등록된 구현이 없습니다.

Tasks

AI Agent

Similar Papers 제목 키워드 기반

From Instructions to Intrinsic Human Values -- A Survey of Alignment Goals for Big Models

2023-08-23 · Jing Yao, Xiaoyuan Yi, Xiting Wang, Jindong Wang 외

Big models, exemplified by Large Language Models (LLMs), are models typically pre-trained on massive data and comprised of enormous parameters, which not only obtain significantly improved performance across diverse task…

Alignment as Jurisprudence

2026-05-08 · Nicholas Caputo arxiv

Jurisprudence, the study of how judges should properly decide cases, and alignment, the science of getting AI models to conform to human values, share a fundamental structure. These seemingly distant fields both seek to …

Superalignment with Dynamic Human Values

2025-03-17 · Florian Mai, David Kaczér, Nicholas Kluge Corrêa, Lucie Flek

Two core challenges of alignment are 1) scalable oversight and 2) accounting for the dynamic nature of human values. While solutions like recursive reward modeling address 1), they do not simultaneously account for 2). W…

Flexible Agent Alignment with Goal Inference from Open-Ended Dialog

2025-08-20 · Rachel Ma, Jingyi Qu, Andreea Bobu, Dylan Hadfield-Menell arxiv

We introduce Open-Universe Assistance Games (OU-AGs), a formal framework extending assistance games to LLM-based agents. Effective assistance requires reasoning over human preferences that are unbounded, underspecified, …

Ethics2vec: aligning automatic agents and human preferences

2025-08-11 · Gianluca Bontempi arxiv

Though intelligent agents are supposed to improve human experience (or make it more efficient), it is hard from a human perspective to grasp the ethical values which are explicitly or implicitly embedded in an agent beha…

Recommendation Systems