Evaluating subgroup disparity using epistemic uncertainty in mammography
As machine learning (ML) continue to be integrated into healthcare systems that affect clinical decision making, new strategies will need to be incorporated in order to effectively detect and evaluate subgroup disparities to ensure accountability and generalizability in clinical workflows. In this paper, we explore how epistemic uncertainty can be used to evaluate disparity in patient demographics (race) and data acquisition (scanner) subgroups for breast density assessment on a dataset of 108,190 mammograms collected from 33 clinical sites. Our results show that even if aggregate performance is comparable, the choice of uncertainty quantification metric can significantly the subgroup level. We hope this analysis can promote further work on how uncertainty can be leveraged to increase transparency of machine learning applications for clinical deployment.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningDecision MakingUncertainty QuantificationSimilar Papers 제목 키워드 기반
Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity
Choosing when to retrain a deployed classifier requires assessing subgroup error rates across the sequence of models used, including periods between updates. We compare complete scheduled, loss-triggered, and subgroup-ga…
Three Applications of Conformal Prediction for Rating Breast Density in Mammography
Breast cancer is the most common cancers and early detection from mammography screening is crucial in improving patient outcomes. Assessing mammographic breast density is clinically important as the denser breasts have h…
Conformal PredictionDeep LearningFairnessPrediction+1Detecting and Monitoring Bias for Subgroups in Breast Cancer Detection AI
Automated mammography screening plays an important role in early breast cancer detection. However, current machine learning models, developed on some training datasets, may exhibit performance degradation and bias when d…
Breast Cancer DetectionEvaluating Simple Debiasing Techniques in RoBERTa-based Hate Speech Detection Models
The hate speech detection task is known to suffer from bias against African American English (AAE) dialect text, due to the annotation bias present in the underlying hate speech datasets used to train these models. This …
Hate Speech DetectionBias and Generalizability of Foundation Models across Datasets in Breast Mammography
Over the past decades, computer-aided diagnosis tools for breast cancer have been developed to enhance screening procedures, yet their clinical adoption remains challenged by data variability and inherent biases. Althoug…
Domain AdaptationFairnessTransfer Learning