paper-with-me

홈 › Papers

Off-Policy Evaluation with Out-of-Sample Guarantees

2023-01-20 · Sofia Ek, Dave Zachariah, Fredrik D. Johansson, Petre Stoica

We consider the problem of evaluating the performance of a decision policy using past observational data. The outcome of a policy is measured in terms of a loss (aka. disutility or negative reward) and the main problem is making valid inferences about its out-of-sample loss when the past data was observed under a different and possibly unknown policy. Using a sample-splitting method, we show that it is possible to draw such inferences with finite-sample coverage guarantees about the entire loss distribution, rather than just its mean. Importantly, the method takes into account model misspecifications of the past policy - including unmeasured confounding. The evaluation method can be used to certify the performance of a policy using observational data under a specified range of credible model assumptions.

📄 PDF Abstract BibTeX arXiv:2301.08649

Code (1)

sofiaek/off-policy-evaluation 공식 구현

Tasks

Off-policy evaluationvalid

Similar Papers 제목 키워드 기반

Stochastic first-order methods for average-reward Markov decision processes

2022-05-11 · Tianjiao Li, Feiyang Wu, Guanghui Lan

We study average-reward Markov decision processes (AMDPs) and develop novel first-order methods with strong theoretical guarantees for both policy optimization and policy evaluation. Compared with intensive research effo…

Policy Gradient Methods

More for Less: Safe Policy Improvement With Stronger Performance Guarantees

2023-05-13 · Patrick Wienhöft, Marnix Suilen, Thiago D. Simão, Clemens Dubslaff 외

In an offline reinforcement learning setting, the safe policy improvement (SPI) problem aims to improve the performance of a behavior policy according to which sample data has been generated. State-of-the-art approaches …

Krylov-Bellman boosting: Super-linear policy evaluation in general state spaces

2022-10-20 · Eric Xia, Martin J. Wainwright

We present and analyze the Krylov-Bellman Boosting (KBB) algorithm for policy evaluation in general state spaces. It alternates between fitting the Bellman residual using non-parametric regression (as in boosting), and e…

A Practical Guide of Off-Policy Evaluation for Bandit Problems

2020-10-23 · Masahiro Kato, Kenshi Abe, Kaito Ariu, Shota Yasui

Off-policy evaluation (OPE) is the problem of estimating the value of a target policy from samples obtained via different policies. Recently, applying OPE methods for bandit problems has garnered attention. For the theor…

Off-policy evaluation

Conformal Off-Policy Prediction in Contextual Bandits

2022-06-09 · Muhammad Faaiz Taufiq, Jean-Francois Ton, Rob Cornish, Yee Whye Teh 외

Most off-policy evaluation methods for contextual bandits have focused on the expected outcome of a policy, which is estimated via methods that at best provide only asymptotic guarantees. However, in many applications, t…

Conformal PredictionMulti-Armed BanditsOff-policy evaluationPrediction