paper-with-me

홈 › Papers

Understanding the Curse of Horizon in Off-Policy Evaluation via Conditional Importance Sampling

2019-10-15 · ICML 2020 1 · Yao Liu, Pierre-Luc Bacon, Emma Brunskill

Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods that leverage the structure of Markov decision processes. We analyze the variance of the most popular approaches through the viewpoint of conditional Monte Carlo. Surprisingly, we find that in finite horizon MDPs there is no strict variance reduction of per-decision importance sampling or stationary importance sampling, comparing with vanilla importance sampling. We then provide sufficient conditions under which the per-decision or stationary estimators will provably reduce the variance over importance sampling with finite horizons. For the asymptotic (in terms of horizon $T$) case, we develop upper and lower bounds on the variance of those estimators which yields sufficient conditions under which there exists an exponential v.s. polynomial gap between the variance of importance sampling and that of the per-decision or stationary estimators. These results help advance our understanding of if and when new types of IS estimators will improve the accuracy of off-policy estimation.

📄 PDF Abstract BibTeX arXiv:1910.06508

Code (0)

등록된 구현이 없습니다.

Tasks

Off-policy evaluationReinforcement Learning

Similar Papers 제목 키워드 기반

Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning

2019-09-12 · Nathan Kallus, Masatoshi Uehara

Off-policy evaluation (OPE) in reinforcement learning is notoriously difficult in long- and infinite-horizon settings due to diminishing overlap between behavior and target policies. In this paper, we study the role of M…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Future-Dependent Value-Based Off-Policy Evaluation in POMDPs

2022-07-26 · NeurIPS 2023 11 · Masatoshi Uehara, Haruka Kiyohara, Andrew Bennett, Victor Chernozhukov 외

We study off-policy evaluation (OPE) for partially observable MDPs (POMDPs) with general function approximation. Existing methods such as sequential importance sampling estimators and fitted-Q evaluation suffer from the …

Off-policy evaluation

Off-policy Evaluation in Infinite-Horizon Reinforcement Learning with Latent Confounders

2020-07-27 · Andrew Bennett, Nathan Kallus, Lihong Li, Ali Mousavi

Off-policy evaluation (OPE) in reinforcement learning is an important problem in settings where experimentation is limited, such as education and healthcare. But, in these very same settings, observed actions are often c…

Off-policy evaluationreinforcement-learningReinforcement Learning (RL)

Black-box Off-policy Estimation for Infinite-Horizon Reinforcement Learning

2020-03-24 · ICLR 2020 1 · Ali Mousavi, Lihong Li, Qiang Liu, Denny Zhou

Off-policy estimation for long-horizon problems is important in many real-life applications such as healthcare and robotics, where high-fidelity simulators may not be available and on-policy evaluation is expensive or im…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Asymptotically Efficient Off-Policy Evaluation for Tabular Reinforcement Learning

2020-01-29 · Ming Yin, Yu-Xiang Wang

We consider the problem of off-policy evaluation for reinforcement learning, where the goal is to estimate the expected reward of a target policy $\pi$ using offline data collected by running a logging policy $\mu$. Stan…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)