Marginalized Operators for Off-policy Reinforcement Learning
In this work, we propose marginalized operators, a new class of off-policy evaluation operators for reinforcement learning. Marginalized operators strictly generalize generic multi-step operators, such as Retrace, as special cases. Marginalized operators also suggest a form of sample-based estimates with potential variance reduction, compared to sample-based estimates of the original multi-step operators. We show that the estimates for marginalized operators can be computed in a scalable way, which also generalizes prior results on marginalized importance sampling as special cases. Finally, we empirically demonstrate that marginalized operators provide performance gains to off-policy evaluation and downstream policy optimization algorithms.
Code (0)
등록된 구현이 없습니다.
Tasks
Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor Representation
Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. Howev…
Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learning+2Practical Marginalized Importance Sampling with the Successor Representation
Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. Howev…
Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learning+2Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling
Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the…
Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Double Reinforcement Learning for Efficient and Robust Off-Policy Evaluation
Off-policy evaluation (OPE) in reinforcement learning allows one to evaluate novel decision policies without needing to conduct exploration, which is often costly or otherwise infeasible. We consider for the first time t…
Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes
Off-policy evaluation (OPE) in reinforcement learning allows one to evaluate novel decision policies without needing to conduct exploration, which is often costly or otherwise infeasible. We consider for the first time t…
Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)