On the Relation between Prediction and Imputation Accuracy under Missing Covariates
Missing covariates in regression or classification problems can prohibit the direct use of advanced tools for further analysis. Recent research has realized an increasing trend towards the usage of modern Machine Learning algorithms for imputation. It originates from their capability of showing favourable prediction accuracy in different learning problems. In this work, we analyze through simulation the interaction between imputation accuracy and prediction accuracy in regression learning problems with missing covariates when Machine Learning based methods for both, imputation and prediction are used. In addition, we explore imputation performance when using statistical inference procedures in prediction settings, such as coverage rates of (valid) prediction intervals. Our analysis is based on empirical datasets provided by the UCI Machine Learning repository and an extensive simulation study.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningImputationPredictionPrediction IntervalsregressionRelationvalidSimilar Papers 제목 키워드 기반
Multiple imputation using chained random forests: a preliminary study based on the empirical distribution of out-of-bag prediction errors
Missing data are common in data analyses in biomedical fields, and imputation methods based on random forests (RF) have become widely accepted, as the RF algorithm can achieve high accuracy without the need for specifica…
ImputationPredictionvalidComparison of Missing Data Imputation Methods using the Framingham Heart study dataset
Cardiovascular disease (CVD) is a class of diseases that involve the heart or blood vessels and according to World Health Organization is the leading cause of death worldwide. EHR data regarding this case, as well as med…
ImputationMissing ValuesImputation techniques on missing values in breast cancer treatment and fertility data
Clinical decision support using data mining techniques offers more intelligent way to reduce the decision error in the last few years. However, clinical datasets often suffer from high missingness, which adversely impact…
ImputationMissing ValuesMultiple-level Point Embedding for Solving Human Trajectory Imputation with Prediction
Sparsity is a common issue in many trajectory datasets, including human mobility data. This issue frequently brings more difficulty to relevant learning tasks, such as trajectory imputation and prediction. Nowadays, litt…
DecoderImputationPredictionGEDI: A Graph-based End-to-end Data Imputation Framework
Data imputation is an effective way to handle missing data, which is common in practical applications. In this study, we propose and test a novel data imputation process that achieve two important goals: (1) preserve the…
Graph structure learningImputationMeta-LearningPrediction