paper-with-me

홈 › Papers

Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

2019-11-15 · Cameron Voloshin, Hoang M. Le, Nan Jiang, Yisong Yue

We offer an experimental benchmark and empirical study for off-policy policy evaluation (OPE) in reinforcement learning, which is a key problem in many safety critical applications. Given the increasing interest in deploying learning-based methods, there has been a flurry of recent proposals for OPE method, leading to a need for standardized empirical analyses. Our work takes a strong focus on diversity of experimental design to enable stress testing of OPE methods. We provide a comprehensive benchmarking suite to study the interplay of different attributes on method performance. We distill the results into a summarized set of guidelines for OPE in practice. Our software package, the Caltech OPE Benchmarking Suite (COBS), is open-sourced and we invite interested researchers to further contribute to the benchmark.

📄 PDF Abstract BibTeX arXiv:1911.06854

Code (3)

clvoloshin/COBS 공식 구현 pytorch
clvoloshin/OPE-tools 공식 구현 tf
pearl-utexas/sope

Tasks

BenchmarkingDiversityExperimental Designreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement Learning

2021-11-29 · Rujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht 외

Reinforcement learning (RL) algorithms are often categorized as either on-policy or off-policy depending on whether they use data from a target policy of interest or from a different behavior policy. In this paper, we st…

Offline RLreinforcement-learningReinforcement Learning (RL)

Beyond the One Step Greedy Approach in Reinforcement Learning

2018-02-10 · Yonathan Efroni, Gal Dalal, Bruno Scherrer, Shie Mannor

The famous Policy Iteration algorithm alternates between policy improvement and policy evaluation. Implementations of this algorithm with several variants of the latter evaluation stage, e.g, $n$-step and trace-based ret…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Beyond the One-Step Greedy Approach in Reinforcement Learning

2018-07-01 · ICML 2018 7 · Yonathan Efroni, Gal Dalal, Bruno Scherrer, Shie Mannor

The famous Policy Iteration algorithm alternates between policy improvement and policy evaluation. Implementations of this algorithm with several variants of the latter evaluation stage, e.g, n-step and trace-based …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Doubly Optimal Policy Evaluation for Reinforcement Learning

2024-10-03 · Shuze Liu, Claire Chen, Shangtong Zhang

Policy evaluation estimates the performance of a policy by (1) collecting data from the environment and (2) processing raw data into a meaningful estimate. Due to the sequential nature of reinforcement learning, any impr…

reinforcement-learningReinforcement Learning

More Efficient Off-Policy Evaluation through Regularized Targeted Learning

2019-12-13 · Aurélien F. Bibaut, Ivana Malenica, Nikos Vlassis, Mark J. Van Der Laan

We study the problem of off-policy evaluation (OPE) in Reinforcement Learning (RL), where the aim is to estimate the performance of a new policy given historical data that may have been generated by a different policy, o…

Causal InferenceOff-policy evaluationReinforcement LearningReinforcement Learning (RL)