paper-with-me

Papers

Quantile-Optimal Policy Learning under Unmeasured Confounding

2025-06-08 · Zhongren Chen, Siyu Chen, Zhengling Qi, Xiaohong Chen, Zhuoran Yang

We study quantile-optimal policy learning where the goal is to find a policy whose reward distribution has the largest $\alpha$-quantile for some $\alpha \in (0, 1)$. We focus on the offline setting whose generating process involves unobserved confounders. Such a problem suffers from three main challenges: (i) nonlinearity of the quantile objective as a functional of the reward distribution, (ii) unobserved confounding issue, and (iii) insufficient coverage of the offline dataset. To address these challenges, we propose a suite of causal-assisted policy learning methods that provably enjoy strong theoretical guarantees under mild conditions. In particular, to address (i) and (ii), using causal inference tools such as instrumental variables and negative controls, we propose to estimate the quantile objectives by solving nonlinear functional integral equations. Then we adopt a minimax estimation approach with nonparametric models to solve these integral equations, and propose to construct conservative policy estimates that address (iii). The final policy is the one that maximizes these pessimistic estimates. In addition, we propose a novel regularized policy learning method that is more amenable to computation. Finally, we prove that the policies learned by these methods are $\tilde{\mathscr{O}}(n^{-1/2})$ quantile-optimal under a mild coverage assumption on the offline dataset. Here, $\tilde{\mathscr{O}}(\cdot)$ omits poly-logarithmic factors. To the best of our knowledge, we propose the first sample-efficient policy learning algorithms for estimating the quantile-optimal policy when there exist unmeasured confounding.

📄 PDF Abstract BibTeX arXiv:2506.07140

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Inference

Methods 이 논문이 사용한 방법론

ADOPT Please enter a description about the method here
Focus 설명 없음
Causal inference Causal inference is the process of drawing a conclusion about a causal connection based on the conditions of the occurrence of an effect. The main difference between causal…

Similar Papers 제목 키워드 기반

Individualized Decision-Making Under Partial Identification: Three Perspectives, Two Optimality Results, and One Paradox

2021-10-21 · Yifan Cui

Unmeasured confounding is a threat to causal inference and gives rise to biased estimates. In this article, we consider the problem of individualized decision-making under partial identification. Firstly, we argue that w…

Causal InferenceDecision Making

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding

2025-05-01 · Yuhan Li, Eugene Han, Yifan Hu, Wenzhuo Zhou 외

This paper addresses the challenge of offline policy learning in reinforcement learning with continuous action spaces when unmeasured confounders are present. While most existing research focuses on policy evaluation wit…

reinforcement-learningReinforcement Learning

Estimating and Improving Dynamic Treatment Regimes With a Time-Varying Instrumental Variable

2021-04-15 · Shuxiao Chen, Bo Zhang

Estimating dynamic treatment regimes (DTRs) from retrospective observational data is challenging as some degree of unmeasured confounding is often expected. In this work, we develop a framework of estimating properly def…

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

2024-12-08 · Shuguang Yu, Shuxing Fang, Ruixin Peng, Zhengling Qi 외

This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data literature, we propose a two-way unmeasured c…

Off-policy evaluation

Proximal Learning for Individualized Treatment Regimes Under Unmeasured Confounding

2021-05-03 · Zhengling Qi, Rui Miao, Xiaoke Zhang

Data-driven individualized decision making has recently received increasing research interests. Most existing methods rely on the assumption of no unmeasured confounding, which unfortunately cannot be ensured in practice…

Causal InferenceDecision Making