A Quantitative Evaluation Framework for Missing Value Imputation Algorithms
We consider the problem of quantitatively evaluating missing value imputation algorithms. Given a dataset with missing values and a choice of several imputation algorithms to fill them in, there is currently no principled way to rank the algorithms using a quantitative metric. We develop a framework based on treating imputation evaluation as a problem of comparing two distributions and show how it can be used to compute quantitative metrics. We present an efficient procedure for applying this framework to practical datasets, demonstrate several metrics derived from the existing literature on comparing distributions, and propose a new metric called Neighborhood-based Dissimilarity Score which is fast to compute and provides similar results. Results are shown on several datasets, metrics, and imputations algorithms.
Code (0)
등록된 구현이 없습니다.
Tasks
ImputationMissing ValuesSimilar Papers 제목 키워드 기반
Multistage Large Segment Imputation Framework Based on Deep Learning and Statistic Metrics
Missing value is a very common and unavoidable problem in sensors, and researchers have made numerous attempts for missing value imputation, particularly in deep learning models. However, for real sensor data, the specif…
ImputationStill More Shades of Null: An Evaluation Suite for Responsible Missing Value Imputation
Data missingness is a practical challenge of sustained interest to the scientific community. In this paper, we present Shades-of-Null, an evaluation suite for responsible missing value imputation. Our work is novel in tw…
FairnessImputationDeep Imputation of Missing Values in Time Series Health Data: A Review with Benchmarking
The imputation of missing values in multivariate time series (MTS) data is critical in ensuring data quality and producing reliable data-driven predictive models. Apart from many statistical approaches, a few recent stud…
BenchmarkingDeep LearningImputationMissing Values+2Comparison of Missing Data Imputation Methods using the Framingham Heart study dataset
Cardiovascular disease (CVD) is a class of diseases that involve the heart or blood vessels and according to World Health Organization is the leading cause of death worldwide. EHR data regarding this case, as well as med…
ImputationMissing ValuesSeeing Through the Clouds: Cloud Gap Imputation with Prithvi Foundation Model
Filling cloudy pixels in multispectral satellite imagery is essential for accurate data analysis and downstream applications, especially for tasks which require time series data. To address this issue, we compare the per…
Generative Adversarial NetworkImputationTime Series