Confidence intervals for AB-test
AB-testing is a very popular technique in web companies since it makes it possible to accurately predict the impact of a modification with the simplicity of a random split across users. One of the critical aspects of an AB-test is its duration and it is important to reliably compute confidence intervals associated with the metric of interest to know when to stop the test. In this paper, we define a clean mathematical framework to model the AB-test process. We then propose three algorithms based on bootstrapping and on the central limit theorem to compute reliable confidence intervals which extend to other metrics than the common probabilities of success. They apply to both absolute and relative increments of the most used comparison metrics, including the number of occurrences of a particular event and a click-through rate implying a ratio.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Confidence Intervals for Testing Disparate Impact in Fair Learning
We provide the asymptotic distribution of the major indexes used in the statistical literature to quantify disparate treatment in machine learning. We aim at promoting the use of confidence intervals when testing the so-…
BIG-bench Machine LearningConfidence Intervals for Performance Estimates in Brain MRI Segmentation
Medical segmentation models are evaluated empirically. As such an evaluation is based on a limited set of example images, it is unavoidably noisy. Beyond a mean performance measure, reporting confidence intervals is thus…
Brain Tumor SegmentationHippocampusImage SegmentationMedical Image Segmentation+4Testing and Confidence Intervals for High Dimensional Proportional Hazards Model
This paper proposes a decorrelation-based approach to test hypotheses and construct confidence intervals for the low dimensional component of high dimensional proportional hazards models. Motivated by the geometric proje…
Model SelectionVocal Bursts Intensity PredictionCross-validation Confidence Intervals for Test Error
This work develops central limit theorems for cross-validation and consistent estimators of its asymptotic variance under weak stability conditions on the learning algorithm. Together, these results provide practical, as…
validConfidence Intervals and Hypothesis Testing for High-Dimensional Statistical Models
Fitting high-dimensional statistical models often requires the use of non-linear parameter estimation procedures. As a consequence, it is generally impossible to obtain an exact characterization of the probability distri…
Diabetes Predictionparameter estimationTwo-sample testingVocal Bursts Intensity Prediction