paper-with-me

홈 › Papers

Performance of weakly-supervised electronic health record-based phenotyping methods in rare-outcome settings

2026-04-10 · Yunjing Hong, Jennifer C. Nelson, Brian D. Williamson arxiv

Accurately identifying patients with specific medical conditions is a key challenge when using clinical data from electronic health records. Our objective was to comprehensively assess when weakly-supervised prediction methods, which use silver-standard labels (proxy measures of the true outcome) rather than gold-standard true labels, perform well in rare-outcome settings like vaccine safety studies. We compared three methods (PheNorm, MAP, and sureLDA) that combine structured features and features derived from clinical text using natural language processing, through an extensive simulation study with data-generating mechanisms ranging from simple to complex, varying outcome rates, and varying degrees of informative silver labels. We also considered using predicted probabilities to design a chart review validation study. No single method dominated the other across all prediction performance metrics. Probability-guided sampling selected a cohort enriched for patients with more mentions of important concepts in chart notes. SureLDA, the most complex of the three algorithms we considered, often performed well in simulations. Performance depended greatly on selected tuning parameters. Care should be taken when using weakly-supervised prediction methods in rare-outcome settings, particularly if the probabilities will be used in downstream analysis, but these methods can work well when silver labels are strong predictors of true outcomes.

📄 PDF Abstract BibTeX arXiv:2604.09913

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Ontology-driven weak supervision for clinical entity classification in electronic health records

2020-08-05 · Jason A. Fries, Ethan Steinberg, Saelig Khattar, Scott L. Fleming 외

In the electronic health record, using clinical notes to identify entities such as disorders and their temporality (e.g. the order of an event relative to a time index) can inform many important analyses. However, creati…

General ClassificationNamed Entity Recognition (NER)Temporal Information ExtractionWeakly Supervised Classification+1

Learning Longitudinal Health Representations from EHR and Wearable Data

2026-01-18 · Yuanyun Zhang, Han Zhou, Li Feng, Yilin Hong 외 arxiv

Foundation models trained on electronic health records show strong performance on many clinical prediction tasks but are limited by sparse and irregular documentation. Wearable devices provide dense continuous physiologi…

Enhancing Phenotype Discovery in Electronic Health Records through Prior Knowledge-Guided Unsupervised Learning

2025-11-03 · Melanie Mayer, Kimberly Lactaoen, Gary E. Weissman, Blanca E. Himes 외 arxiv

Objectives: Unsupervised learning with electronic health record (EHR) data has shown promise for phenotype discovery, but approaches typically disregard existing clinical information, limiting interpretability. We operat…

Clinical Knowledge

Unsupervised Pseudo-Labeling for Extractive Summarization on Electronic Health Records

2018-11-20 · Xiangan Liu, Keyang Xu, Pengtao Xie, Eric Xing

Extractive summarization is very useful for physicians to better manage and digest Electronic Health Records (EHRs). However, the training of a supervised model requires disease-specific medical background and is thus ve…

Extractive Summarization

A Semi-supervised Approach for De-identification of Swedish Clinical Text

2020-05-01 · LREC 2020 5 · Hanna Berg, Hercules Dalianis

An abundance of electronic health records (EHR) is produced every day within healthcare. The records possess valuable information for research and future improvement of healthcare. Multiple efforts have been done to prot…

De-identification