paper-with-me

홈 › Papers

Parenting: Safe Reinforcement Learning from Human Input

2019-02-18 · Christopher Frye, Ilya Feige

Autonomous agents trained via reinforcement learning present numerous safety concerns: reward hacking, negative side effects, and unsafe exploration, among others. In the context of near-future autonomous agents, operating in environments where humans understand the existing dangers, human involvement in the learning process has proved a promising approach to AI Safety. Here we demonstrate that a precise framework for learning from human input, loosely inspired by the way humans parent children, solves a broad class of safety problems in this context. We show that our Parenting algorithm solves these problems in the relevant AI Safety gridworlds of Leike et al. (2017), that an agent can learn to outperform its parent as it "matures", and that policies learnt through Parenting are generalisable to new environments.

📄 PDF Abstract BibTeX arXiv:1902.06766

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement Learning

Similar Papers 제목 키워드 기반

Parenting: Optimizing Knowledge Selection of Retrieval-Augmented Language Models with Parameter Decoupling and Tailored Tuning

2024-10-14 · Yongxin Xu, Ruizhe Zhang, Xinke Jiang, Yujie Feng 외

Retrieval-Augmented Generation (RAG) offers an effective solution to the issues faced by Large Language Models (LLMs) in hallucination generation and knowledge obsolescence by incorporating externally retrieved knowledge…

HallucinationRAGRetrievalRetrieval-augmented Generation

Theories of Parenting and their Application to Artificial Intelligence

2019-03-14 · Sky Croeser, Peter Eckersley

As machine learning (ML) systems have advanced, they have acquired more power over humans' lives, and questions about what values are embedded in them have become more complex and fraught. It is conceivable that in the c…

Ethics

PARENTing via Model-Agnostic Reinforcement Learning to Correct Pathological Behaviors in Data-to-Text Generation

2020-10-21 · INLG (ACL) 2020 12 · Clément Rebuffel, Laure Soulier, Geoffrey Scoutheeten, Patrick Gallinari

In language generation models conditioned by structured data, the classical training via maximum likelihood almost always leads models to pick up on dataset divergence (i.e., hallucinations or omissions), and to incorpor…

Data-to-Text Generationreinforcement-learningReinforcement Learning (RL)Text Generation

Supertrust foundational alignment: mutual trust must replace permanent control for safe superintelligence

2024-07-29 · James M. Mazzu

It's widely expected that humanity will someday create AI systems vastly more intelligent than us, leading to the unsolved alignment problem of "how to control superintelligence." However, this commonly expressed problem…

Interacting safely with cyclists using Hamilton-Jacobi reachability and reinforcement learning

2026-02-20 · Aarati Andrea Noronha, Jean Oh arxiv

In this paper, we present a framework for enabling autonomous vehicles to interact with cyclists in a manner that balances safety and optimality. The approach integrates Hamilton-Jacobi reachability analysis with deep Q-…

Reinforcement LearningAutonomous Vehicles