paper-with-me

홈 › Papers

Mind the Performance Gap: Examining Dataset Shift During Prospective Validation

2021-07-23 · Erkin Ötleş, Jeeheh Oh, Benjamin Li, Michelle Bochinski, Hyeon Joo, Justin Ortwine, Erica Shenoy, Laraine Washer, Vincent B. Young, Krishna Rao, Jenna Wiens

Once integrated into clinical care, patient risk stratification models may perform worse compared to their retrospective performance. To date, it is widely accepted that performance will degrade over time due to changes in care processes and patient populations. However, the extent to which this occurs is poorly understood, in part because few researchers report prospective validation performance. In this study, we compare the 2020-2021 ('20-'21) prospective performance of a patient risk stratification model for predicting healthcare-associated infections to a 2019-2020 ('19-'20) retrospective validation of the same model. We define the difference in retrospective and prospective performance as the performance gap. We estimate how i) "temporal shift", i.e., changes in clinical workflows and patient populations, and ii) "infrastructure shift", i.e., changes in access, extraction and transformation of data, both contribute to the performance gap. Applied prospectively to 26,864 hospital encounters during a twelve-month period from July 2020 to June 2021, the model achieved an area under the receiver operating characteristic curve (AUROC) of 0.767 (95% confidence interval (CI): 0.737, 0.801) and a Brier score of 0.189 (95% CI: 0.186, 0.191). Prospective performance decreased slightly compared to '19-'20 retrospective performance, in which the model achieved an AUROC of 0.778 (95% CI: 0.744, 0.815) and a Brier score of 0.163 (95% CI: 0.161, 0.165). The resulting performance gap was primarily due to infrastructure shift and not temporal shift. So long as we continue to develop and validate models using data stored in large research data warehouses, we must consider differences in how and when data are accessed, measure how these differences may affect prospective performance, and work to mitigate those differences.

📄 PDF Abstract BibTeX arXiv:2107.13964

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Brain Resection Multimodal Image Registration (ReMIND2Reg) 2025 Challenge

2025-08-13 · Reuben Dorent, Laura Rigolo, Colin P. Galvin, Junyu Chen 외 arxiv

Accurate intraoperative image guidance is critical for achieving maximal safe resection in brain tumor surgery, yet neuronavigation systems based on preoperative MRI lose accuracy during the procedure due to brain shift.…

Image Registration

Explicit Modelling of Theory of Mind for Belief Prediction in Nonverbal Social Interactions

2024-07-09 · Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling

We propose MToMnet - a Theory of Mind (ToM) neural network for predicting beliefs and their dynamics during human social interactions from multimodal input. ToM is key for effective nonverbal human communication and coll…

Re-Ranking

Mind the Gap: Bridging Prior Shift in Realistic Few-Shot Crop-Type Classification

2025-11-20 · Joana Reuss, Ekaterina Gikalo, Marco Körner arxiv

Real-world agricultural distributions often suffer from severe class imbalance, typically following a long-tailed distribution. Labeled datasets for crop-type classification are inherently scarce and remain costly to obt…

Few-Shot Learning

MindShift: Leveraging Large Language Models for Mental-States-Based Problematic Smartphone Use Intervention

2023-09-28 · Ruolan Wu, Chun Yu, Xiaole Pan, Yujia Liu 외

Problematic smartphone use negatively affects physical and mental health. Despite the wide range of prior research, existing persuasive techniques are not flexible enough to provide dynamic persuasion content based on us…

Persuasion Strategies

Using LLMs to Infer Non-Binary COVID-19 Sentiments of Chinese Micro-bloggers

2025-01-09 · Jerry Chongyi Hu, Mohammed Shahid Modi, Boleslaw K. Szymanski

Studying public sentiment during crises is crucial for understanding how opinions and sentiments shift, resulting in polarized societies. We study Weibo, the most popular microblogging site in China, using posts made dur…

Language ModelingLanguage ModellingLarge Language ModelSentiment Analysis