paper-with-me

홈 › Papers

Collapsibility of Performance Metrics in Clinical Predictive AI

2026-08-31 · João Matos, Ben Van Calster, Richard D. Riley, Paula Dhiman, Gary S. Collins arxiv

Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup specific values. Objective: To examine the collapsibility properties of commonly reported performance metrics in predictive AI, with a focus on the area under the receiver operating characteristic curve (AUC, also known as c-statistic). Methods: We investigate the collapsibility of 15 performance metrics, either by expressing each metric as a linear combination of its stratum specific values or, where non-collapsible, by providing a counterexample inspired by Simpson's paradox as a formal disproof. Results: Five performance metrics (AUC, calibration intercept, calibration slope, expected calibration error, and Nagelkerke R^2) are shown to be non-collapsible, and ten (O:E ratio, logloss, Brier score, accuracy, F1-score, true positive rate, true negative rate, positive predictive value, negative predictive value, and net benefit) are shown to be collapsible. The AUC is shown to be non-collapsible because it decomposes into within- and cross-group AUC terms when subpopulations coexist, such that its overall value may fall outside the range of subgroup specific AUCs. Conclusions: Non-collapsibility of performance metrics has important consequences for reporting, model appraisal, and fairness evaluation. It can generate spurious differences between subgroup and overall performance, which may mislead fairness evaluations. Explicitly acknowledging and reporting the collapsibility properties of performance metrics improves both the interpretability and transparency of fairness assessments.

📄 PDF Abstract BibTeX arXiv:2608.30568

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Estimate Collapsibility of Causal Effects in Completed Partial DAGs via Strong d-Convex Hulls

2026-06-08 · Yuxin Deng, Yi Sun, Zhiming Li, Huaxiong Liu arxiv

This paper proposes a collapsible method for estimating causal effects that maintains the estimator's consistency before and after marginalization over some variables in completed partially directed acyclic graphs (CPDAG…

The role of the geometric mean in case-control studies

2022-07-19 · Amanda Coston, Edward H. Kennedy

Historically used in settings where the outcome is rare or data collection is expensive, outcome-dependent sampling is relevant to many modern settings where data is readily available for a biased sample of the target po…

Evaluation of Predictive Data Mining Algorithms in Erythemato-Squamous Disease Diagnosis

2015-01-03 · Kwetishe Danjuma, Adenike O. Osofisan

A lot of time is spent searching for the most performing data mining algorithms applied in clinical diagnosis. The study set out to identify the most performing predictive data mining algorithms applied in the diagnosis …

Novel Techniques to Assess Predictive Systems and Reduce Their Alarm Burden

2021-02-10 · Jonathan A. Handler, Craig F. Feied, Michael T. Gillam

Machine prediction algorithms (e.g., binary classifiers) often are adopted on the basis of claimed performance using classic metrics such as sensitivity and predictive value. However, classifier performance depends heavi…

Explainable artificial intelligence model to predict acute critical illness from electronic health records

2019-12-03 · Simon Meyer Lauritsen, Mads Kristensen, Mathias Vassard Olsen, Morten Skaarup Larsen 외

We developed an explainable artificial intelligence (AI) early warning score (xAI-EWS) system for early detection of acute critical illness. While maintaining a high predictive performance, our system explains to the cli…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)PredictionSpecificity+1