paper-with-me

홈 › Papers

Towards Optimal Off-Policy Evaluation for Reinforcement Learning with Marginalized Importance Sampling

2019-06-08 · NeurIPS 2019 12 · Tengyang Xie, Yifei Ma, Yu-Xiang Wang

Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the historical data obtained by different behavior policies -- under the model of nonstationary episodic Markov Decision Processes (MDP) with a long horizon and a large action space. Existing importance sampling (IS) methods often suffer from large variance that depends exponentially on the RL horizon $H$. To solve this problem, we consider a marginalized importance sampling (MIS) estimator that recursively estimates the state marginal distribution for the target policy at every step. MIS achieves a mean-squared error of $$ \frac{1}{n} \sum\nolimits_{t=1}^H\mathbb{E}_{\mu}\left[\frac{d_t^\pi(s_t)^2}{d_t^\mu(s_t)^2} \mathrm{Var}_{\mu}\left[\frac{\pi_t(a_t|s_t)}{\mu_t(a_t|s_t)}\big( V_{t+1}^\pi(s_{t+1}) + r_t\big) \middle| s_t\right]\right] + \tilde{O}(n^{-1.5}) $$ where $\mu$ and $\pi$ are the logging and target policies, $d_t^{\mu}(s_t)$ and $d_t^{\pi}(s_t)$ are the marginal distribution of the state at $t$th step, $H$ is the horizon, $n$ is the sample size and $V_{t+1}^\pi$ is the value function of the MDP under $\pi$. The result matches the Cramer-Rao lower bound in \citet{jiang2016doubly} up to a multiplicative factor of $H$. To the best of our knowledge, this is the first OPE estimation error bound with a polynomial dependence on $H$. Besides theory, we show empirical superiority of our method in time-varying, partially observable, and long-horizon RL environments.

📄 PDF Abstract BibTeX arXiv:1906.03393

Code (0)

등록된 구현이 없습니다.

Tasks

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Marginalized Operators for Off-policy Reinforcement Learning

2022-03-30 · Yunhao Tang, Mark Rowland, Rémi Munos, Michal Valko

In this work, we propose marginalized operators, a new class of off-policy evaluation operators for reinforcement learning. Marginalized operators strictly generalize generic multi-step operators, such as Retrace, as spe…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor Representation

2021-06-12 · Scott Fujimoto, David Meger, Doina Precup

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. Howev…

Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learning+2

Importance Weighted Actor-Critic for Optimal Conservative Offline Reinforcement Learning

2023-01-30 · NeurIPS 2023 11 · Hanlin Zhu, Paria Rashidinejad, Jiantao Jiao

We propose A-Crab (Actor-Critic Regularized by Average Bellman error), a new practical algorithm for offline reinforcement learning (RL) in complex environments with insufficient data coverage. Our algorithm combines the…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Practical Marginalized Importance Sampling with the Successor Representation

2021-01-01 · Scott Fujimoto, David Meger, Doina Precup

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. Howev…

Deep Reinforcement LearningMuJoCoOff-policy evaluationreinforcement-learning+2

Minimax Weight and Q-Function Learning for Off-Policy Evaluation

2019-10-28 · ICML 2020 1 · Masatoshi Uehara, Jiawei Huang, Nan Jiang

We provide theoretical investigations into off-policy evaluation in reinforcement learning using function approximators for (marginalized) importance weights and value functions. Our contributions include: (1) A new esti…

Off-policy evaluationReinforcement LearningReinforcement Learning (RL)