paper-with-me

홈 › Papers

Consequences of Misaligned AI

2021-02-07 · NeurIPS 2020 12 · Simon Zhuang, Dylan Hadfield-Menell

AI systems often rely on two key components: a specified goal or reward function and an optimization algorithm to compute the optimal behavior for that goal. This approach is intended to provide value for a principal: the user on whose behalf the agent acts. The objectives given to these agents often refer to a partial specification of the principal's goals. We consider the cost of this incompleteness by analyzing a model of a principal and an agent in a resource constrained world where the $L$ attributes of the state correspond to different sources of utility for the principal. We assume that the reward function given to the agent only has support on $J < L$ attributes. The contributions of our paper are as follows: 1) we propose a novel model of an incomplete principal-agent problem from artificial intelligence; 2) we provide necessary and sufficient conditions under which indefinitely optimizing for any incomplete proxy objective leads to arbitrarily low overall utility; and 3) we show how modifying the setup to allow reward functions that reference the full state or allowing the principal to update the proxy objective over time can lead to higher utility solutions. The results in this paper argue that we should view the design of reward functions as an interactive and dynamic process and identifies a theoretical scenario where some degree of interactivity is desirable.

📄 PDF Abstract BibTeX arXiv:2102.03896

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Preemptive Detection and Correction of Misaligned Actions in LLM Agents

2024-07-16 · Haishuo Fang, Xiaodan Zhu, Iryna Gurevych

Deploying LLM-based agents in real-life applications often faces a critical challenge: the misalignment between agents' behavior and user intent. Such misalignment may lead agents to unintentionally execute critical acti…

Action DetectionDecision Making

Score Design for Multi-Criteria Incentivization

2024-10-08 · Anmol Kabra, Mina Karzand, Tosca Lechner, Nathan Srebro 외

We present a framework for designing scores to summarize performance metrics. Our design has two multi-criteria objectives: (1) improving on scores should improve all performance metrics, and (2) achieving pareto-optimal…

When Language Models Lose Their Mind: The Consequences of Brain Misalignment

2026-03-24 · Gabriele Merlin, Mariya Toneva arxiv

While brain-aligned large language models (LLMs) have garnered attention for their potential as cognitive models and for potential for enhanced safety and trustworthiness in AI, the role of this brain alignment for lingu…

RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation

2025-01-15 · Kaiqu Liang, Haimin Hu, Ryan Liu, Thomas L. Griffiths 외

Generative AI systems like foundation models (FMs) must align well with human values to ensure their behavior is helpful and trustworthy. While Reinforcement Learning from Human Feedback (RLHF) has shown promise for opti…

Technology Readiness Levels for AI & ML

2020-06-21 · Alexander Lavin, Gregory Renard

The development and deployment of machine learning systems can be executed easily with modern tools, but the process is typically rushed and means-to-an-end. The lack of diligence can lead to technical debt, scope creep …

BIG-bench Machine Learning