paper-with-me

홈 › Papers

From Weighting to Modeling: A Nonparametric Estimator for Off-Policy Evaluation

2026-03-10 · Rong J. B. Zhu arxiv

We study off-policy evaluation in the setting of contextual bandits, where we aim to evaluate a new policy using historical data that consists of contexts, actions and received rewards. This historical data typically does not faithfully represent action distribution of the new policy accurately. A common approach, inverse probability weighting (IPW), adjusts for these discrepancies in action distributions. However, this method often suffers from high variance due to the probability being in the denominator. The doubly robust (DR) estimator reduces variance through modeling reward but does not directly address variance from IPW. In this work, we address the limitation of IPW by proposing a Nonparametric Weighting (NW) approach that constructs weights using a nonparametric model. Our NW approach achieves low bias like IPW but typically exhibits significantly lower variance. To further reduce variance, we incorporate reward predictions -- similar to the DR technique -- resulting in the Model-assisted Nonparametric Weighting (MNW) approach. The MNW approach yields accurate value estimates by explicitly modeling and mitigating bias from reward modeling, without aiming to guarantee the standard doubly robust property. Extensive empirical comparisons show that our approaches consistently outperform existing techniques, achieving lower variance in value estimation while maintaining low bias.

📄 PDF Abstract BibTeX arXiv:2603.09436

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Nonparametric inverse probability weighted estimators based on the highly adaptive lasso

2020-05-22 · Ashkan Ertefaie, Nima S. Hejazi, Mark J. Van Der Laan

Inverse probability weighted estimators are the oldest and potentially most commonly used class of procedures for the estimation of causal effects. By adjusting for selection biases via a weighting mechanism, these proce…

On Well-posedness and Minimax Optimal Rates of Nonparametric Q-function Estimation in Off-policy Evaluation

2022-01-17 · Xiaohong Chen, Zhengling Qi

We study the off-policy evaluation (OPE) problem in an infinite-horizon Markov decision process with continuous states and actions. We recast the $Q$-function estimation into a special form of the nonparametric instrumen…

Off-policy evaluation

Off-Policy Evaluation for Large Action Spaces via Conjunct Effect Modeling

2023-05-14 · Yuta Saito, Qingyang Ren, Thorsten Joachims

We study off-policy evaluation (OPE) of contextual bandit policies for large discrete action spaces where conventional importance-weighting approaches suffer from excessive variance. To circumvent this variance issue, we…

Off-policy evaluation

Off-Policy Evaluation and Learning for External Validity under a Covariate Shift

2020-02-26 · NeurIPS 2020 12 · Masahiro Kato, Masatoshi Uehara, Shota Yasui

We consider evaluating and training a new policy for the evaluation data by using the historical data obtained from a different policy. The goal of off-policy evaluation (OPE) is to estimate the expected reward of a new …

Off-policy evaluation

Off-Policy Evaluation via Adaptive Weighting with Data from Contextual Bandits

2021-06-03 · Ruohan Zhan, Vitor Hadad, David A. Hirshberg, Susan Athey

It has become increasingly common for data to be collected adaptively, for example using contextual bandits. Historical data of this type can be used to evaluate other treatment assignment policies to guide future innova…

Multi-Armed BanditsOff-policy evaluation