paper-with-me

홈 › Papers

Marginalized Operators for Off-policy Reinforcement Learning

2022-03-30 · Yunhao Tang, Mark Rowland, Rémi Munos, Michal Valko

In this work, we propose marginalized operators, a new class of off-policy evaluation operators for reinforcement learning. Marginalized operators strictly generalize generic multi-step operators, such as Retrace, as special cases. Marginalized operators also suggest a form of sample-based estimates with potential variance reduction, compared to sample-based estimates of the original multi-step operators. We show that the estimates for marginalized operators can be computed in a scalable way, which also generalizes prior results on marginalized importance sampling as special cases. Finally, we empirically demonstrate that marginalized operators provide performance gains to off-policy evaluation and downstream policy optimization algorithms.

📄 PDF Abstract BibTeX arXiv:2203.16177

Code (0)

등록된 구현이 없습니다.

Tasks

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Retrace Retrace is an off-policy Q-value estimation algorithm which has guaranteed convergence for a target and behaviour policy $\left(\pi, \beta\right)$. With off-policy rollout for…

Similar Papers 제목 키워드 기반

A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor Representation

2021-06-12 · Scott Fujimoto, David Meger, Doina Precup

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. Howev…

Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learning+2

Practical Marginalized Importance Sampling with the Successor Representation

2021-01-01 · Scott Fujimoto, David Meger, Doina Precup

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. Howev…

Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learning+2

Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling

2019-06-08 · NeurIPS 2019 12 · Tengyang Xie, Yifei Ma, Yu-Xiang Wang

Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Double Reinforcement Learning for Efficient and Robust Off-Policy Evaluation

2020-01-01 · ICML 2020 1 · Nathan Kallus, Masatoshi Uehara

Off-policy evaluation (OPE) in reinforcement learning allows one to evaluate novel decision policies without needing to conduct exploration, which is often costly or otherwise infeasible. We consider for the first time t…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes

2019-08-22 · Nathan Kallus, Masatoshi Uehara

Off-policy evaluation (OPE) in reinforcement learning allows one to evaluate novel decision policies without needing to conduct exploration, which is often costly or otherwise infeasible. We consider for the first time t…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)