Off-policy Evaluation in Infinite-Horizon Reinforcement Learning with Latent Confounders
Off-policy evaluation (OPE) in reinforcement learning is an important problem in settings where experimentation is limited, such as education and healthcare. But, in these very same settings, observed actions are often confounded by unobserved variables making OPE even more difficult. We study an OPE problem in an infinite-horizon, ergodic Markov decision process with unobserved confounders, where states and actions can act as proxies for the unobserved confounders. We show how, given only a latent variable model for states and actions, policy value can be identified from off-policy data. Our method involves two stages. In the first, we show how to use proxies to estimate stationary distribution ratios, extending recent work on breaking the curse of horizon to the confounded setting. In the second, we show optimal balancing can be combined with such learned ratios to obtain policy value while avoiding direct modeling of reward functions. We establish theoretical guarantees of consistency, and benchmark our method empirically.
Code (0)
등록된 구현이 없습니다.
Tasks
Off-policy evaluationreinforcement-learningReinforcement Learning (RL)Similar Papers 제목 키워드 기반
Confounding-Robust Policy Evaluation in Infinite-Horizon Reinforcement Learning
Off-policy evaluation of sequential decision policies from observational data is necessary in applications of batch reinforcement learning such as education and healthcare. In such settings, however, unobserved variables…
Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1On the Sample Complexity of Vanilla Model-Based Offline Reinforcement Learning with Dependent Samples
Offline reinforcement learning (offline RL) considers problems where learning is performed using only previously collected samples and is helpful for the settings in which collecting new data is costly or risky. In model…
Offline RLOff-policy evaluationreinforcement-learningReinforcement Learning+1Beyond dynamic programming
In this paper, we present Score-life programming, a novel theoretical approach for solving reinforcement learning problems. In contrast with classical dynamic programming-based methods, our method can search over non-sta…
reinforcement-learningReinforcement LearningInfinite Time Horizon Safety of Bayesian Neural Networks
Bayesian neural networks (BNNs) place distributions over the weights of a neural network to model uncertainty in the data and the network's prediction. We consider the problem of verifying safety when running a Bayesian …
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Safe ExplorationStatistical Inference of the Value Function for Reinforcement Learning in Infinite Horizon Settings
Reinforcement learning is a general technique that allows an agent to learn an optimal policy and interact with an environment in sequential decision making problems. The goodness of a policy is measured by its value fun…
Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1