Computational Efficient Approximations of the Concordance Probability in a Big Data Setting
Performance measurement is an essential task once a statistical model is created. The Area Under the receiving operating characteristics Curve (AUC) is the most popular measure for evaluating the quality of a binary classifier. In this case, AUC is equal to the concordance probability, a frequently used measure to evaluate the discriminatory power of the model. Contrary to AUC, the concordance probability can also be extended to the situation with a continuous response variable. Due to the staggering size of data sets nowadays, determining this discriminatory measure requires a tremendous amount of costly computations and is hence immensely time consuming, certainly in case of a continuous response variable. Therefore, we propose two estimation methods that calculate the concordance probability in a fast and accurate way and that can be applied to both the discrete and continuous setting. Extensive simulation studies show the excellent performance and fast computing times of both estimators. Finally, experiments on two real-life data sets confirm the conclusions of the artificial simulations.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Concordance probability in a big data setting: application in non-life insurance
The concordance probability or C-index is a popular measure to capture the discriminatory ability of a regression model. In this article, the definition of this measure is adapted to the specific needs of the frequency a…
regressionDisappointment concordance and duet expectiles
We introduce an axiom of disappointment-concordance (disco) aversion for a preference relation over acts in an Anscombe-Aumann setting. This axiom means that the decision maker, facing the sum of two acts, dislikes the s…
Fast Fair Regression via Efficient Approximations of Mutual Information
Most work in algorithmic fairness to date has focused on discrete outcomes, such as deciding whether to grant someone a loan or not. In these classification settings, group fairness criteria such as independence, separat…
Computational EfficiencyFairnessregressionScalable and consistent embedding of probability measures into Hilbert spaces via measure quantization
This paper is focused on statistical learning from data that come as probability measures. In this setting, popular approaches consist in embedding such data into a Hilbert space with either Linearized Optimal Transport …
QuantizationInductive Coherence
While probability theory is normally applied to external environments, there has been some recent interest in probabilistic modeling of the outputs of computations that are too expensive to run. Since mathematical logic …
NegationSentence