paper-with-me

홈 › Papers

Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

2022-09-18 · Zuyue Fu, Zhengling Qi, Zhaoran Wang, Zhuoran Yang, Yanxun Xu, Michael R. Kosorok

We study the offline reinforcement learning (RL) in the face of unmeasured confounders. Due to the lack of online interaction with the environment, offline RL is facing the following two significant challenges: (i) the agent may be confounded by the unobserved state variables; (ii) the offline data collected a prior does not provide sufficient coverage for the environment. To tackle the above challenges, we study the policy learning in the confounded MDPs with the aid of instrumental variables. Specifically, we first establish value function (VF)-based and marginalized importance sampling (MIS)-based identification results for the expected total reward in the confounded MDPs. Then by leveraging pessimism and our identification results, we propose various policy learning methods with the finite-sample suboptimality guarantee of finding the optimal in-class policy under minimal data coverage and modeling assumptions. Lastly, our extensive theoretical investigations and one numerical study motivated by the kidney transplantation demonstrate the promising performance of the proposed methods.

📄 PDF Abstract BibTeX arXiv:2209.08666

Code (0)

등록된 구현이 없습니다.

Tasks

Offline RLreinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning

2021-02-19 · Luofeng Liao, Zuyue Fu, Zhuoran Yang, Yixin Wang 외

In offline reinforcement learning (RL) an optimal policy is learned solely from a priori collected observational data. However, in observational data, actions are often confounded by unobserved variables. Instrumental va…

Offline RLreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

Statistical Estimation of Confounded Linear MDPs: An Instrumental Variable Approach

2022-09-12 · Miao Lu, Wenhao Yang, Liangyu Zhang, Zhihua Zhang

In an Markov decision process (MDP), unobservable confounders may exist and have impacts on the data generating process, so that the classic off-policy evaluation (OPE) estimators may fail to identify the true value func…

Off-policy evaluation

An Instrumental Variable Approach to Confounded Off-Policy Evaluation

2022-12-29 · Yang Xu, Jin Zhu, Chengchun Shi, Shikai Luo 외

Off-policy evaluation (OPE) is a method for estimating the return of a target policy using some pre-collected observational data generated by a potentially different behavior policy. In some cases, there may be unmeasure…

Decision MakingOff-policy evaluation

Confounded Causal Imitation Learning with Instrumental Variables

2025-07-23 · Yan Zeng, Shenglan Nie, Feng Xie, Libo Huang 외 arxiv

Imitation learning from demonstrations usually suffers from the confounding effects of unmeasured variables (i.e., unmeasured confounders) on the states and actions. If ignoring them, a biased estimation of the policy wo…

Learning Decision Policies with Instrumental Variables through Double Machine Learning

2024-05-14 · Daqian Shao, Ashkan Soleymani, Francesco Quinzan, Marta Kwiatkowska

A common issue in learning decision-making policies in data-rich settings is spurious correlations in the offline dataset, which can be caused by hidden confounders. Instrumental variable (IV) regression, which utilises …

Decision Makingregression