High-Throughput Machine Learning from Electronic Health Records
The widespread digitization of patient data via electronic health records (EHRs) has created an unprecedented opportunity to use machine learning algorithms to better predict disease risk at the patient level. Although predictive models have previously been constructed for a few important diseases, such as breast cancer and myocardial infarction, we currently know very little about how accurately the risk for most diseases or events can be predicted, and how far in advance. Machine learning algorithms use training data rather than preprogrammed rules to make predictions and are well suited for the complex task of disease prediction. Although there are thousands of conditions and illnesses patients can encounter, no prior research simultaneously predicts risks for thousands of diagnosis codes and thereby establishes a comprehensive patient risk profile. Here we show that such pandiagnostic prediction is possible with a high level of performance across diagnosis codes. For the tasks of predicting diagnosis risks both 1 and 6 months in advance, we achieve average areas under the receiver operating characteristic curve (AUCs) of 0.803 and 0.758, respectively, across thousands of prediction tasks. Finally, our research contributes a new clinical prediction dataset in which researchers can explore how well a diagnosis can be predicted and what health factors are most useful for prediction. For the first time, we can get a much more complete picture of how well risks for thousands of different diagnosis codes can be predicted.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningDisease PredictionPredictionVocal Bursts Intensity PredictionSimilar Papers 제목 키워드 기반
High Throughput Phenotyping of Physician Notes with Large Language and Hybrid NLP Models
Deep phenotyping is the detailed description of patient signs and symptoms using concepts from an ontology. The deep phenotyping of the numerous physician notes in electronic health records requires high throughput metho…
Language ModelingLanguage ModellingLarge Language ModelPrivacy Preserving Machine Learning for Electronic Health Records using Federated Learning and Differential Privacy
An Electronic Health Record (EHR) is an electronic database used by healthcare providers to store patients' medical records which may include diagnoses, treatments, costs, and other personal information. Machine learning…
Federated LearningPrivacy PreservingA Large Language Model Outperforms Other Computational Approaches to the High-Throughput Phenotyping of Physician Notes
High-throughput phenotyping, the automated mapping of patient signs and symptoms to standardized ontology concepts, is essential to gaining value from electronic health records (EHR) in the support of precision medicine.…
Language ModelingLanguage ModellingLarge Language ModelDeveloping A Visual-Interactive Interface for Electronic Health Record Labeling: An Explainable Machine Learning Approach
Labeling a large number of electronic health records is expensive and time consuming, and having a labeling assistant tool can significantly reduce medical experts' workload. Nevertheless, to gain the experts' trust, the…
High-throughput relation extraction algorithm development associating knowledge articles and electronic health records
Objective: Medical relations are the core components of medical knowledge graphs that are needed for healthcare artificial intelligence. However, the requirement of expert annotation by conventional algorithm development…
ArticlesKnowledge GraphsRelationRelation Extraction+2