paper-with-me

홈 › Papers

Policy Shaping: Integrating Human Feedback with Reinforcement Learning

2013-12-01 · NeurIPS 2013 12 · Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L. Isbell, Andrea L. Thomaz

A long term goal of Interactive Reinforcement Learning is to incorporate non-expert human feedback to solve complex tasks. State-of-the-art methods have approached this problem by mapping human information to reward and value signals to indicate preferences and then iterating over them to compute the necessary control policy. In this paper we argue for an alternate, more effective characterization of human feedback: Policy Shaping. We introduce Advise, a Bayesian approach that attempts to maximize the information gained from human feedback by utilizing it as direct labels on the policy. We compare Advise to state-of-the-art approaches and highlight scenarios where it outperforms them and importantly is robust to infrequent and inconsistent human feedback.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Influencing Reinforcement Learning through Natural Language Guidance

2021-04-04 · Tasmia Tasrin, Md Sultan Al Nahian, Habarakadage Perera, Brent Harrison

Interactive reinforcement learning agents use human feedback or instruction to help them learn in complex environments. Often, this feedback comes in the form of a discrete signal that is either positive or negative. Whi…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Learning Efficient Dialogue Policy from Demonstrations through Shaping

2020-07-01 · ACL 2020 6 · Huimin Wang, Baolin Peng, Kam-Fai Wong

Training a task-oriented dialogue agent with reinforcement learning is prohibitively expensive since it requires a large volume of interactions with users. Human demonstrations can be used to accelerate learning progress…

Domain Adaptation

Interactive Reinforcement Learning for Table Balancing Robot

2021-08-01 · ACL (splurobonlp) 2021 8 · Haein Jeon, Yewon Kim, Bo-Yeong Kang

With the development of robotics, the use of robots in daily life is increasing, which has led to the need for anyone to easily train robots to improve robot use. Interactive reinforcement learning(IARL) is a method for …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Deep Reinforcement Learningreinforcement-learning+5

Navigating Noisy Feedback: Enhancing Reinforcement Learning with Error-Prone Language Models

2024-10-22 · MuHan Lin, Shuyang Shi, Yue Guo, Behdad Chalaki 외

The correct specification of reward models is a well-known challenge in reinforcement learning. Hand-crafted reward functions often lead to inefficient or suboptimal policies and may not be aligned with user values. Rein…

HallucinationLanguage ModelingLanguage ModellingLarge Language Model+2

Environment Shaping in Reinforcement Learning using State Abstraction

2020-06-23 · Parameswaran Kamalaruban, Rati Devidze, Volkan Cevher, Adish Singla

One of the central challenges faced by a reinforcement learning (RL) agent is to effectively learn a (near-)optimal policy in environments with large state spaces having sparse and noisy feedback signals. In real-world a…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)