paper-with-me

홈 › Papers

Tipping the Balance: Impact of Class Imbalance Correction on the Performance of Clinical Risk Prediction Models

2026-02-27 · Amalie Koch Andersen, Hadi Mehdizavareh, Arijit Khan, Tobias Becher, Simone Britsch, Markward Britsch, Morten Bøttcher, Simon Winther, Palle Duun Rohde, Morten Hasselstrøm Jensen, Simon Lebech Cichosz arxiv

Objective: ML-based clinical risk prediction models are increasingly used to support decision-making in healthcare. While class-imbalance correction techniques are commonly applied to improve model performance in settings with rare outcomes, their impact on probabilistic calibration remains insufficiently understood. This study evaluated the effect of widely used resampling strategies on both discrimination and calibration across real-world clinical prediction tasks. Methods: Ten clinical datasets spanning diverse medical domains and including 605,842 patients were analyzed. Multiple machine-learning model families, including linear models and several non-linear approaches, were evaluated. Models were trained on the original data and under three commonly used 1:1 class-imbalance correction strategies (SMOTE, RUS, ROS). Performance was assessed on held-out data using discrimination and calibration metrics. Results: Across all datasets and model families, resampling had no positive impact on predictive performance. Changes in the Receiver Operating Characteristic Area Under Curve (ROC-AUC) relative to models trained on the original data were small and inconsistent (ROS: -0.002, p<0.05; RUS: -0.004, p>0.05; SMOTE: -0.01, p<0.05), with no resampling strategy demonstrating a systematic improvement. In contrast, resampling in general degraded the calibration performance. Models trained using imbalance correction exhibited higher Brier scores (0.029 to 0.080, p<0.05), reflecting poorer probabilistic accuracy, and marked deviations in calibration intercept and slope, indicating systematic distortions of predicted risk despite preserved rank-based performance. Conclusion: In a diverse set of real-world clinical prediction tasks, commonly used class-imbalance correction techniques did not provide generalizable improvements in discrimination and were associated with degraded calibration.

📄 PDF Abstract BibTeX arXiv:2603.00208

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Addressing Fairness, Bias and Class Imbalance in Machine Learning: the FBI-loss

2021-05-13 · Elisa Ferrari, Davide Bacciu

Resilience to class imbalance and confounding biases, together with the assurance of fairness guarantees are highly desirable properties of autonomous decision-making systems with real-life impact. Many different targete…

BIG-bench Machine LearningDecision MakingFairness

Let the Fuzzy Rule Speak: Enhancing In-context Learning Debiasing with Interpretability

2024-12-26 · Ruixi Lin, Yang You

One of the potential failures of large language models (LLMs) is their imbalanced class performances in text classification tasks. With in-context learning (ICL), LLMs yields good accuracy for some classes but low accura…

In-Context LearningMulti Class Text Classificationtext-classificationText Classification

Twice Class Bias Correction for Imbalanced Semi-Supervised Learning

2023-12-27 · Lan Li, Bowen Tao, Lu Han, De-Chuan Zhan 외

Differing from traditional semi-supervised learning, class-imbalanced semi-supervised learning presents two distinct challenges: (1) The imbalanced distribution of training samples leads to model bias towards certain cla…

Bayes Imbalance Impact Index: A Measure of Class Imbalanced Dataset for Classification Problem

2019-01-29 · Yang Lu, Yiu-ming Cheung, Yuan Yan Tang

Recent studies have shown that imbalance ratio is not the only cause of the performance loss of a classifier in imbalanced data classification. In fact, other data factors, such as small disjuncts, noises and overlapping…

General Classification

Phased Progressive Learning with Coupling-Regulation-Imbalance Loss for Imbalanced Data Classification

2022-05-24 · Liang Xu, Yi Cheng, Fan Zhang, Bingxuan Wu 외

Deep convolutional neural networks often perform poorly when faced with datasets that suffer from quantity imbalances and classification difficulties. Despite advances in the field, existing two-stage approaches still ex…

Classificationimbalanced classificationRepresentation Learning