Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning
We offer an experimental benchmark and empirical study for off-policy policy evaluation (OPE) in reinforcement learning, which is a key problem in many safety critical applications. Given the increasing interest in deploying learning-based methods, there has been a flurry of recent proposals for OPE method, leading to a need for standardized empirical analyses. Our work takes a strong focus on diversity of experimental design to enable stress testing of OPE methods. We provide a comprehensive benchmarking suite to study the interplay of different attributes on method performance. We distill the results into a summarized set of guidelines for OPE in practice. Our software package, the Caltech OPE Benchmarking Suite (COBS), is open-sourced and we invite interested researchers to further contribute to the benchmark.
Code (3)
Tasks
BenchmarkingDiversityExperimental Designreinforcement-learningReinforcement LearningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement Learning
Reinforcement learning (RL) algorithms are often categorized as either on-policy or off-policy depending on whether they use data from a target policy of interest or from a different behavior policy. In this paper, we st…
Offline RLreinforcement-learningReinforcement Learning (RL)Beyond the One Step Greedy Approach in Reinforcement Learning
The famous Policy Iteration algorithm alternates between policy improvement and policy evaluation. Implementations of this algorithm with several variants of the latter evaluation stage, e.g, $n$-step and trace-based ret…
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Beyond the One-Step Greedy Approach in Reinforcement Learning
The famous Policy Iteration algorithm alternates between policy improvement and policy evaluation. Implementations of this algorithm with several variants of the latter evaluation stage, e.g, n-step and trace-based …
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Doubly Optimal Policy Evaluation for Reinforcement Learning
Policy evaluation estimates the performance of a policy by (1) collecting data from the environment and (2) processing raw data into a meaningful estimate. Due to the sequential nature of reinforcement learning, any impr…
reinforcement-learningReinforcement LearningMore Efficient Off-Policy Evaluation through Regularized Targeted Learning
We study the problem of off-policy evaluation (OPE) in Reinforcement Learning (RL), where the aim is to estimate the performance of a new policy given historical data that may have been generated by a different policy, o…
Causal InferenceOff-policy evaluationReinforcement LearningReinforcement Learning (RL)