paper-with-me

홈 › Papers

Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning

2021-02-19 · Luofeng Liao, Zuyue Fu, Zhuoran Yang, Yixin Wang, Mladen Kolar, Zhaoran Wang

In offline reinforcement learning (RL) an optimal policy is learned solely from a priori collected observational data. However, in observational data, actions are often confounded by unobserved variables. Instrumental variables (IVs), in the context of RL, are the variables whose influence on the state variables is all mediated by the action. When a valid instrument is present, we can recover the confounded transition dynamics through observational data. We study a confounded Markov decision process where the transition dynamics admit an additive nonlinear functional form. Using IVs, we derive a conditional moment restriction through which we can identify transition dynamics based on observational data. We propose a provably efficient IV-aided Value Iteration (IVVI) algorithm based on a primal-dual reformulation of the conditional moment restriction. To our knowledge, this is the first provably efficient algorithm for instrument-aided offline RL.

📄 PDF Abstract BibTeX arXiv:2102.09907

Code (0)

등록된 구현이 없습니다.

Tasks

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)valid

Similar Papers 제목 키워드 기반

Casual Social Media Use among the Youth: Effects on Online and Offline Political Participation

2023-12-14 · Mehdi Barati

Background: Previous studies suggest that social media use among the youth is correlated with online and offline political participation. There is also a mixed and inconclusive debate on whether more online political par…

Least Squares Policy Iteration with Instrumental Variables vs. Direct Policy Search: Comparison Against Optimal Benchmarks Using Energy Storage

2014-01-04 · Warren R. Scott, Warren B. Powell, Somayeh Moazehi

This paper studies approximate policy iteration (API) methods which use least-squares Bellman error minimization for policy evaluation. We address several of its enhancements, namely, Bellman error minimization using ins…

On Instrumental Variable Regression for Deep Offline Policy Evaluation

2021-05-21 · Yutian Chen, Liyuan Xu, Caglar Gulcehre, Tom Le Paine 외

We show that the popular reinforcement learning (RL) strategy of estimating the state-action value (Q-function) by minimizing the mean squared Bellman error leads to a regression problem with confounding, the inputs and …

regressionReinforcement Learning (RL)

Personalized Pricing with Invalid Instrumental Variables: Identification, Estimation, and Policy Learning

2023-02-24 · Rui Miao, Zhengling Qi, Cong Shi, Lin Lin

Pricing based on individual customer characteristics is widely used to maximize sellers' revenues. This work studies offline personalized pricing under endogeneity using an instrumental variable approach. Standard instru…

Causal InferenceEconometrics

Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

2022-09-18 · Zuyue Fu, Zhengling Qi, Zhaoran Wang, Zhuoran Yang 외

We study the offline reinforcement learning (RL) in the face of unmeasured confounders. Due to the lack of online interaction with the environment, offline RL is facing the following two significant challenges: (i) the a…

Offline RLreinforcement-learningReinforcement Learning (RL)