paper-with-me

홈 › Papers

Off-Policy Evaluation and Learning for External Validity under a Covariate Shift

2020-02-26 · NeurIPS 2020 12 · Masahiro Kato, Masatoshi Uehara, Shota Yasui

We consider evaluating and training a new policy for the evaluation data by using the historical data obtained from a different policy. The goal of off-policy evaluation (OPE) is to estimate the expected reward of a new policy over the evaluation data, and that of off-policy learning (OPL) is to find a new policy that maximizes the expected reward over the evaluation data. Although the standard OPE and OPL assume the same distribution of covariate between the historical and evaluation data, a covariate shift often exists, i.e., the distribution of the covariate of the historical data is different from that of the evaluation data. In this paper, we derive the efficiency bound of OPE under a covariate shift. Then, we propose doubly robust and efficient estimators for OPE and OPL under a covariate shift by using a nonparametric estimator of the density ratio between the historical and evaluation data distributions. We also discuss other possible estimators and compare their theoretical properties. Finally, we confirm the effectiveness of the proposed estimators through experiments.

📄 PDF Abstract BibTeX arXiv:2002.11642

Code (1)

MasaKat0/OPE_CS 공식 구현

Tasks

Off-policy evaluation

Similar Papers 제목 키워드 기반

Externally Valid Policy Evaluation Combining Trial and Observational Data

2023-10-23 · Sofia Ek, Dave Zachariah

Randomized trials are widely considered as the gold standard for evaluating the effects of decision policies. Trial data is, however, drawn from a population which may differ from the intended target population and this …

valid

Robustness, Heterogeneous Treatment Effects and Covariate Shifts

2021-12-16 · Pietro Emilio Spini

This paper studies the robustness of estimated policy effects to changes in the distribution of covariates. Robustness to covariate shifts is important, for example, when evaluating the external validity of quasi-experim…

Policy Learning under Biased Sample Selection

2023-04-23 · Lihua Lei, Roshni Sahoo, Stefan Wager

Practitioners often use data from a randomized controlled trial to learn a treatment assignment policy that can be deployed on a target population. A recurring concern in doing so is that, even if the randomized trial wa…

Selecting Experimental Sites for External Validity

2024-05-21 · Michael Gechter, Keisuke Hirano, Jean Lee, Mahreen Mahmud 외

Policy decisions often depend on evidence generated elsewhere. We take a Bayesian decision-theoretic approach to choosing where to experiment to optimize external validity. We frame external validity through a policy len…

Externally Valid Selection of Experimental Sites via the k-Median Problem

2024-08-17 · José Luis Montiel Olea, Brenda Prallon, Chen Qiu, Jörg Stoye 외

We present a decision-theoretic justification for viewing the question of how to best choose where to experiment in order to optimize external validity as a k-median (clustering) problem, a popular problem in computer sc…

valid