paper-with-me

Papers

Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework

2023-09-23 · Wenzhuo Zhou, Yuhan Li, Ruoqing Zhu, Annie Qu

We study high-confidence off-policy evaluation in the context of infinite-horizon Markov decision processes, where the objective is to establish a confidence interval (CI) for the target policy value using only offline data pre-collected from unknown behavior policies. This task faces two primary challenges: providing a comprehensive and rigorous error quantification in CI estimation, and addressing the distributional shift that results from discrepancies between the distribution induced by the target policy and the offline data-generating process. Motivated by an innovative unified error analysis, we jointly quantify the two sources of estimation errors: the misspecification error on modeling marginalized importance weights and the statistical uncertainty due to sampling, within a single interval. This unified framework reveals a previously hidden tradeoff between the errors, which undermines the tightness of the CI. Relying on a carefully designed discriminator function, the proposed estimator achieves a dual purpose: breaking the curse of the tradeoff to attain the tightest possible CI, and adapting the CI to ensure robustness against distributional shifts. Our method is applicable to time-dependent data without assuming any weak dependence conditions via leveraging a local supermartingale/martingale structure. Theoretically, we show that our algorithm is sample-efficient, error-robust, and provably convergent even in non-linear function approximation settings. The numerical performance of the proposed method is examined in synthetic datasets and an OhioT1DM mobile health study.

📄 PDF Abstract BibTeX arXiv:2309.13278

Code (0)

등록된 구현이 없습니다.

Tasks

Off-policy evaluation

Similar Papers 제목 키워드 기반

Conformal Prediction Beyond the Horizon: Distribution-Free Inference for Policy Evaluation

2025-10-29 · Feichen Gan, Youcun Lu, Yingying Zhang, Yukun Liu arxiv

Reliable uncertainty quantification is crucial for reinforcement learning (RL) in high-stakes settings. We propose a unified conformal prediction framework for infinite-horizon policy evaluation that constructs distribut…

Reinforcement Learning

The $s$-value: evaluating stability with respect to distributional shifts

2021-05-07 · Suyash Gupta, Dominik Rothenhäusler

Common statistical measures of uncertainty such as $p$-values and confidence intervals quantify the uncertainty due to sampling, that is, the uncertainty due to not observing the full population. However, sampling is not…

Distributionally Robust Policy Evaluation under General Covariate Shift in Contextual Bandits

2024-01-21 · Yihong Guo, Hao liu, Yisong Yue, Anqi Liu

We introduce a distributionally robust approach that enhances the reliability of offline policy evaluation in contextual bandits under general covariate shifts. Our method aims to deliver robust policy evaluation results…

Multi-Armed Banditsregression

The s-value: evaluating stability with respect to distributional shifts

2023-09-21 · NeurIPS 2023 11

Common statistical measures of uncertainty such as $p$-values and confidence intervals quantify the uncertainty due to sampling, that is, the uncertainty due to not observing the full population. However, sampling is not…

Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes

2024-03-29 · Andrew Bennett, Nathan Kallus, Miruna Oprescu, Wen Sun 외

We study the evaluation of a policy under best- and worst-case perturbations to a Markov decision process (MDP), using transition observations from the original MDP, whether they are generated under the same or a differe…

Off-policy evaluation