Safe Policy Learning from Observations
In this paper, we consider the problem of learning a policy by observing numerous non-expert agents. Our goal is to extract a policy that, with high-confidence, acts better than the agents' average performance. Such a setting is important for real-world problems where expert data is scarce but non-expert data can easily be obtained, e.g. by crowdsourcing. Our approach is to pose this problem as safe policy improvement in reinforcement learning. First, we evaluate an average behavior policy and approximate its value function. Then, we develop a stochastic policy improvement algorithm that safely improves the average behavior. The primary advantages of our approach, termed Rerouted Behavior Improvement (RBI), over other safe learning methods are its stability in the presence of value estimation errors and the elimination of a policy search process. We demonstrate these advantages in the Taxi grid-world domain and in four games from the Atari learning environment.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Safe non-smooth black-box optimization with application to policy search
For safety-critical black-box optimization tasks, observations of the constraints and the objective are often noisy and available only for the feasible points. We propose an approach based on log barriers to find a local…
Verifiable Foundation Models for Robot Safety
Deploying foundation models for robot control raises a central challenge: the expressive power that enables rich, multimodal perception also makes these models opaque and difficult to analyze formally, rendering them int…
Collision AvoidanceState-Wise Safe Reinforcement Learning With Pixel Observations
In the context of safe exploration, Reinforcement Learning (RL) has long grappled with the challenges of balancing the tradeoff between maximizing rewards and minimizing safety violations, particularly in complex environ…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Exploration+1Safe Reinforcement Learning in Tensor Reproducing Kernel Hilbert Space
This paper delves into the problem of safe reinforcement learning (RL) in a partially observable environment with the aim of achieving safe-reachability objectives. In traditional partially observable Markov decision pro…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe Reinforcement LearningSafeAPT: Safe Simulation-to-Real Robot Learning using Diverse Policies Learned in Simulation
The framework of Simulation-to-real learning, i.e, learning policies in simulation and transferring those policies to the real world is one of the most promising approaches towards data-efficient learning in robotics. Ho…
Bayesian Optimization