paper-with-me

홈 › Papers

Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap

2026-07-16 · Olivier Jeunen arxiv

Online controlled experiments are the gold standard for hypothesis testing in online platforms. Notwithstanding their ubiquity, they are notoriously expensive to run, and issues of variance hamper statistical power in assessing treatment effects. While standard variance reduction techniques leverage model-based control variates to reduce outcome noise, they remain agnostic to potential structural relationships between competing policies. In this work, we identify a critical inefficiency in the standard A/B-testing protocol: when a treatment and control policy agree on an action, the resulting outcome contributes noise but no signal regarding the treatment effect -- unnecessarily inflating confidence intervals. We propose a novel experimental protocol that exploits this policy overlap to accelerate experimentation. The key insight is to frame the randomised treatment assignment mechanism as a meta-policy, and leverage $Δ$-Off-Policy Estimation methods to obtain unbiased estimates for average treatment effects. We prove analytically that our approach recovers standard A/B-testing practices in the general case, but that its variance scales with the divergence between policies rather than raw outcome variance. Hence, we dominate the standard Difference-in-Means estimator whenever policies have common support, and the improvement is strict whenever the overlap region contributes non-zero residual variance. Empirical results corroborate these theoretical insights -- holding promise for significant impact on the real-world evaluation of recommender systems, information retrieval pipelines, and large language model interfaces.

📄 PDF Abstract BibTeX arXiv:2607.14604

Code (0)

등록된 구현이 없습니다.

Tasks

Information Retrieval

Similar Papers 제목 키워드 기반

Counterfactual Invariance to Spurious Correlations: Why and How to Pass Stress Tests

2021-05-31 · NeurIPS 2021 12 · Victor Veitch, Alexander D'Amour, Steve Yadlowsky, Jacob Eisenstein

Informally, a 'spurious correlation' is the dependence of a model on some aspect of the input data that an analyst thinks shouldn't matter. In machine learning, these have a know-it-when-you-see-it character; e.g., chang…

Causal Inferencecounterfactualtext-classificationText Classification

Counterfactual Explanations for Deep Two-Sample Testing

2026-05-29 · Wei-Cheng Lai, Marco Simnacher, Christoph Lippert arxiv

Two-sample testing is a fundamental tool for detecting distributional differences across scientific domains, but classical tests (including kernel-based tests) can be ineffective on high-dimensional structured data such …

Two-sample testing

Reconsidering Generative Objectives For Counterfactual Reasoning

2020-12-01 · NeurIPS 2020 12 · Danni Lu, Chenyang Tao, Junya Chen, Fan Li 외

There has been recent interest in exploring generative goals for counterfactual reasoning, such as individualized treatment effect (ITE) estimation. However, existing solutions often fail to address issues that are uniqu…

Causal InferencecounterfactualCounterfactual ReasoningRepresentation Learning

Exogenous Matching: Learning Good Proposals for Tractable Counterfactual Estimation

2024-10-17 · Yikang Chen, Dehui Du, Lili Tian

We propose an importance sampling method for tractable and efficient estimation of counterfactual expressions in general settings, named Exogenous Matching. By minimizing a common upper bound of counterfactual estimators…

counterfactual

Density-Guided Robust Counterfactual Explanations on Tabular Data under Model Multiplicity

2026-05-29 · Jun Tan, Qing Guo, Zicheng Xu, Jinglin Li 외 arxiv

Counterfactual explanations (CEs) are essential for actionable recourse, yet their reliability is often compromised in low-density regions, where classifiers exhibit high variance. Unlike existing methods that rely on ex…