paper-with-me

홈 › Papers

Masking criteria for selecting an imputation model

2025-11-13 · Yanjiao Yang, Daniel Suen, Yen-Chi Chen arxiv

The masking-one-out (MOO) procedure, masking an observed entry and comparing it versus its imputed values, is a very common procedure for comparing imputation models. We study the optimum of this procedure and generalize it to a missing data assumption and establish the corresponding semi-parametric efficiency theory. However, MOO is a measure of prediction accuracy, which is not ideal for evaluating an imputation model. To address this issue, we introduce three modified MOO criteria, based on rank transformation, energy distance, and likelihood principle, that allow us to select an imputation model that properly account for the stochastic nature of data. The likelihood approach further enables an elegant framework of learning an imputation model from the data and we derive its statistical and computational learning theories as well as consistency of BIC model selection. We also show how MOO is related to the missing-at-random assumption. Finally, we introduce the prediction-imputation diagram, a two-dimensional diagram visually comparing both the prediction and imputation utilities for various imputation models.

📄 PDF Abstract BibTeX arXiv:2511.10048

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unveiling the Secrets: How Masking Strategies Shape Time Series Imputation

2024-05-26 · Linglong Qian, Yiyuan Yang, Wenjie Du, Jun Wang 외

Time series imputation is a critical challenge in data mining, particularly in domains like healthcare and environmental monitoring, where missing data can compromise analytical outcomes. This study investigates the infl…

ImputationTime Series

Diffusion Models for Tabular Data Imputation and Synthetic Data Generation

2024-07-02 · Mario Villaizán-Vallelado, Matteo Salvatori, Carlos Segura, Ioannis Arapakis

Data imputation and data generation have important applications for many domains, like healthcare and finance, where incomplete or missing data can hinder accurate analysis and decision-making. Diffusion models have emer…

DecoderDenoisingImputationSynthetic Data Generation

Representation Learning for Wearable-Based Applications in the Case of Missing Data

2024-01-08 · Janosch Jungo, Yutong Xiang, Shkurta Gashi, Christian Holz

Wearable devices continuously collect sensor data and use it to infer an individual's behavior, such as sleep, physical activity, and emotions. Despite the significant interest and advancements in this field, modeling mu…

ImputationRepresentation LearningSelf-Supervised Learning

Unmasking Trees for Tabular Data

2024-07-08 · Calvin Mccarter

Despite much work on advanced deep learning and generative modeling techniques for tabular data generation and imputation, traditional methods have continued to win on imputation benchmarks. We herein present UnmaskingTr…

Density EstimationImputationIn-Context LearningTabular Data Generation

To Predict or Not To Predict? Proportionally Masked Autoencoders for Tabular Data Imputation

2024-12-26 · Jungkyu Kim, Kibok Lee, Taeyoung Park

Masked autoencoders (MAEs) have recently demonstrated effectiveness in tabular data imputation. However, due to the inherent heterogeneity of tabular data, the uniform random masking strategy commonly used in MAEs can di…

Imputation