paper-with-me

Papers

The Missing Indicator Method: From Low to High Dimensions

2022-11-16 · Mike Van Ness, Tomas M. Bosschieter, Roberto Halpin-Gregorio, Madeleine Udell

Missing data is common in applied data science, particularly for tabular data sets found in healthcare, social sciences, and natural sciences. Most supervised learning methods only work on complete data, thus requiring preprocessing such as missing value imputation to work on incomplete data sets. However, imputation alone does not encode useful information about the missing values themselves. For data sets with informative missing patterns, the Missing Indicator Method (MIM), which adds indicator variables to indicate the missing pattern, can be used in conjunction with imputation to improve model performance. While commonly used in data science, MIM is surprisingly understudied from an empirical and especially theoretical perspective. In this paper, we show empirically and theoretically that MIM improves performance for informative missing values, and we prove that MIM does not hurt linear models asymptotically for uninformative missing values. Additionally, we find that for high-dimensional data sets with many uninformative indicators, MIM can induce model overfitting and thus test performance. To address this issue, we introduce Selective MIM (SMIM), a novel MIM extension that adds missing indicators only for features that have informative missing patterns. We show empirically that SMIM performs at least as well as MIM in general, and improves MIM for high-dimensional data. Lastly, to demonstrate the utility of MIM on real-world data science tasks, we demonstrate the effectiveness of MIM and SMIM on clinical tasks generated from the MIMIC-III database of electronic health records.

📄 PDF Abstract BibTeX arXiv:2211.09259

Code (1)

mvanness354/missing_indicator_method 공식 구현

Tasks

ImputationMissing ValuesVocal Bursts Intensity Prediction

Methods 이 논문이 사용한 방법론

Test 설명 없음
MIM 설명 없음

Similar Papers 제목 키워드 기반

NeuMiss networks: differentiable programming for supervised learning with missing values

2020-07-03 · Marine Le Morvan, Julie Josse, Thomas Moreau, Erwan Scornet 외

The presence of missing values makes supervised learning much more challenging. Indeed, previous work has shown that even when the response is a linear function of the complete data, the optimal predictor is a complex fu…

ImputationMissing Values

NeuMiss networks: differentiable programming for supervised learning with missing values.

2020-12-01 · NeurIPS 2020 12 · Marine Le Morvan, Julie Josses, Thomas Moreau, Erwan Scornet 외

The presence of missing values makes supervised learning much more challenging. Indeed, previous work has shown that even when the response is a linear function of the complete data, the optimal predictor is a complex fu…

ImputationMissing Values

No imputation without representation

2022-06-28 · Oliver Urs Lenz, Daniel Peralta, Chris Cornelis

By filling in missing values in datasets, imputation allows these datasets to be used with algorithms that cannot handle missing values by themselves. However, missing values may in principle contribute useful informatio…

AttributeImputationMissing Values

Interpretable Generalized Additive Models for Datasets with Missing Values

2024-12-03 · Hayden McTavish, Jon Donnelly, Margo Seltzer, Cynthia Rudin

Many important datasets contain samples that are missing one or more feature values. Maintaining the interpretability of machine learning models in the presence of such missing data is challenging. Singly or multiply imp…

Additive modelsImputationMissing Values

Cultural Diversity and Its Impact on Governance

2021-12-15 · Tomáš Evan, Vladimír Holý

Hofstede's six cultural dimensions make it possible to measure the culture of countries but are criticized for assuming the homogeneity of each country. In this paper, we propose two measures based on Hofstede's cultural…

Cultural Vocal Bursts Intensity PredictionDiversity