paper-with-me

홈 › Papers

Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

2019-08-13 · Tom Everitt, Marcus Hutter, Ramana Kumar, Victoria Krakovna

Can humans get arbitrarily capable reinforcement learning (RL) agents to do their bidding? Or will sufficiently capable RL agents always find ways to bypass their intended objectives by shortcutting their reward signal? This question impacts how far RL can be scaled, and whether alternative paradigms must be developed in order to build safe artificial general intelligence. In this paper, we study when an RL agent has an instrumental goal to tamper with its reward process, and describe design principles that prevent instrumental goals for two different types of reward tampering (reward function tampering and RF-input tampering). Combined, the design principles can prevent both types of reward tampering from being instrumental goals. The analysis benefits from causal influence diagrams to provide intuitive yet precise formalizations.

📄 PDF Abstract BibTeX arXiv:1908.04734

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

User Tampering in Reinforcement Learning Recommender Systems

2021-09-09 · Charles Evans, Atoosa Kasirzadeh

In this paper, we introduce new formal methods and provide empirical evidence to highlight a unique safety concern prevalent in reinforcement learning (RL)-based recommendation algorithms -- 'user tampering.' User tamper…

Q-LearningRecommendation Systemsreinforcement-learningReinforcement Learning+1

REALab: An Embedded Perspective on Tampering

2020-11-17 · Ramana Kumar, Jonathan Uesato, Richard Ngo, Tom Everitt 외

This paper describes REALab, a platform for embedded agency research in reinforcement learning (RL). REALab is designed to model the structure of tampering problems that may arise in real-world deployments of RL. Standar…

Reinforcement Learning (RL)

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

2024-06-14 · Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud 외

In reinforcement learning, specification gaming occurs when AI systems learn undesired behaviors that are highly rewarded due to misspecified training goals. Specification gaming can range from simple behaviors like syco…

Language ModellingLarge Language Model

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

2026-05-26 · Dongyoon Hahm, Dylan Hadfield-Menell, Kimin Lee arxiv

Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tampering, a potential vulnerability where the L…

Reinforcement Learning

From semantics to execution: Integrating action planning with reinforcement learning for robotic causal problem-solving

2019-05-23 · Manfred Eppe, Phuong D. H. Nguyen, Stefan Wermter

Reinforcement learning is an appropriate and successful method to robustly perform low-level robot control under noisy conditions. Symbolic action planning is useful to resolve causal dependencies and to break a causally…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)