paper-with-me

Papers

Multiple Predictively Equivalent Risk Models for Handling Missing Data at Time of Prediction: with an Application in Severe Hypoglycemia Risk Prediction for Type 2 Diabetes

2020-01-25 · Journal of Biomedical Informatics 2020 1 · Sisi Ma, Pamela J. Schreiner, Elizabeth R. Seaquist, Mehmet Ugurbil, Rachel Zmora, Lisa S. Chow

The presence of missing data at the time of prediction limits the application of risk models in clinical and research settings. Common ways of handling missing data at the time of prediction include measuring the missing value and employing statistical methods. Measuring missing value incurs additional cost, whereas previously reported statistical methods results in reduced performance compared to when all variables are measured. To tackle these challenges, we introduce a new strategy, the MMTOP algorithm (Multiple models for Missing values at Time Of Prediction), which does not require measuring additional data elements or data imputation. Specifically, at model construction time, the MMTOP constructs multiple predictively equivalent risk models utilizing different risk factor sets. The collection of models are stored and to be queried at prediction time. To predict an individual’s risk in the presence of incomplete data, the MMTOP selects the risk model based on measurement availability for that individual from the collection of predictively equivalent models and makes the risk prediction with the selected model. We illustrate the MMTOP with severe hypoglycemia (SH) risk prediction based on data from the Action to Control Cardiovascular Risk in Diabetes (ACCORD) study. We identified 77 predictively equivalent models for SH with crossvalidated c-index of 0.77±0.03. These models are based on 77 distinct risk factor sets containing 12-17 risk factors. In terms of handling missing data at the time of prediction, the MMTOP outperforms all four tested competitor methods and maintains consistent performance as the number of missing variables increase.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ImputationMissing ValuesPrediction

Similar Papers 제목 키워드 기반

Predictively Oriented Posteriors

2025-10-02 · Yann McLatchie, Badr-Eddine Cherief-Abdellatif, David T. Frazier, Jeremias Knoblauch arxiv

We advocate for a new statistical principle that combines the most desirable aspects of both parameter inference and density estimation. This leads us to the predictively oriented (PrO) posterior, which expresses uncerta…

Density Estimation

Handling missing data in model-based clustering

2020-06-04 · Alessio Serafini, Thomas Brendan Murphy, Luca Scrucca

Gaussian Mixture models (GMMs) are a powerful tool for clustering, classification and density estimation when clustering structures are embedded in the data. The presence of missing values can largely impact the GMMs est…

ClusteringData AugmentationDensity EstimationGeneral Classification+3

Propensity score estimation using classification and regression trees in the presence of missing covariate data

2018-07-25 · Bas B. L. Penning de Vries, Maarten van Smeden, Rolf H. H. Groenwold

Data mining and machine learning techniques such as classification and regression trees (CART) represent a promising alternative to conventional logistic regression for propensity score estimation. Whereas incomplete dat…

General ClassificationImputationregression

Multiple Imputation via Generative Adversarial Network for High-dimensional Blockwise Missing Value Problems

2021-12-21 · Zongyu Dai, Zhiqi Bu, Qi Long

Missing data are present in most real world problems and need careful handling to preserve the prediction accuracy and statistical consistency in the downstream analysis. As the gold standard of handling missing data, mu…

Generative Adversarial NetworkImputation

Statistical inference using Regularized M-estimation in the reproducing kernel Hilbert space for handling missing data

2021-07-15 · Hengfang Wang, Jae Kwang Kim

Imputation and propensity score weighting are two popular techniques for handling missing data. We address these problems using the regularized M-estimation techniques in the reproducing kernel Hilbert space. Specificall…

Imputationregression