Deeply-Debiased Off-Policy Interval Estimation
Off-policy evaluation learns a target policy's value with a historical dataset generated by a different behavior policy. In addition to a point estimate, many applications would benefit significantly from having a confidence interval (CI) that quantifies the uncertainty of the point estimate. In this paper, we propose a novel deeply-debiasing procedure to construct an efficient, robust, and flexible CI on a target policy's value. Our method is justified by theoretical results and numerical experiments. A Python implementation of the proposed procedure is available at https://github.com/RunzheStat/D2OPE.
Code (1)
Tasks
Off-policy evaluationSimilar Papers 제목 키워드 기반
ScoreMatchingRiesz: Score Matching for Debiased Machine Learning and Policy Path Estimation
We propose ScoreMatchingRiesz, a family of Riesz representer estimators based on score matching. The Riesz representer is a key nuisance component in debiased machine learning, enabling $\sqrt{n}$-consistent and asymptot…
Automatic doubly robust inference for linear functionals via calibrated debiased machine learning
In causal inference, many estimands of interest can be expressed as a linear functional of the outcome regression function; this includes, for example, average causal effects of static, dynamic and stochastic interventio…
Causal InferenceregressionDouble/Debiased Machine Learning for Dynamic Treatment Effects via g-Estimation
We consider the estimation of treatment effects in settings when multiple treatments are assigned over time and treatments can have a causal effect on future outcomes or the state of the treated unit. We propose an exten…
BIG-bench Machine LearningModel SelectionOff-policy evaluationTriple/Debiased Lasso for Statistical Inference of Conditional Average Treatment Effects
This study investigates the estimation and the statistical inference about Conditional Average Treatment Effects (CATEs), which have garnered attention as a metric representing individualized causal effects. In our data-…
regressionUncertainty Quantification for Demand Prediction in Contextual Dynamic Pricing
Data-driven sequential decision has found a wide range of applications in modern operations management, such as dynamic pricing, inventory control, and assortment optimization. Most existing research on data-driven seque…
Assortment OptimizationManagementUncertainty Quantificationvalid