paper-with-me

홈 › Papers

Increasing Missingness to Reduce Bias: Richardson-SGD with Missing Data

2026-05-19 · Ferdinand Genans, Erwan Scornet arxiv

Stochastic gradient methods are central to modern large-scale learning, but their use with incomplete covariates remains delicate since imputation schemes generally introduce systematic gradient biases, as shown for linear models. In this work, we prove that all parametric models exhibit similar gradient bias for various imputation procedures and characterize exactly the dependence on the missingness ratio vector $p$, with $O(\|p\|)$ as the leading term. We exploit this analysis to propose a simple debiasing procedure for stochastic gradient descent (SGD) with missing values based on Richardson extrapolation, which leverages the exact expression of the gradient bias. The key idea is to \emph{deliberately add missingness}: from an already incomplete observation, we generate a further-thinned version at a higher, controlled missingness level, and combine the two resulting stochastic gradients to cancel the leading bias term. We prove that one Richardson step reduces the gradient bias from $O(\|p\|)$ to $O(\|p\|^2)$ under several missingness scenarios. Our proposed method is computationally efficient, model-agnostic and applies to any parametric loss whose stochastic gradient can be computed after imputation. Furthermore, when missing indicators are independent, the population gradient bias is a multilinear polynomial in $p$ and depends only on population gradient errors induced by declaring a single coordinate missing. In this case, our method generalizes to a multi-step Richardson procedure which recursively cancels higher-order terms. Empirically, Richardson debiasing improves optimization and estimation across several generalized linear models and combines positively with widely used imputation procedures such as MICE. These results suggest that, somewhat counter-intuitively, adding controlled missingness on top of existing missing data can make stochastic learning from incomplete data more accurate.

📄 PDF Abstract BibTeX arXiv:2605.19641

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Missingness Bias Calibration in Feature Attribution Explanations

2026-03-05 · Shailesh Sridhar, Anton Xue, Eric Wong arxiv

Popular explanation methods often produce unreliable feature importance scores due to missingness bias, a systematic distortion that arises when models are probed with ablated, out-of-distribution inputs. Existing soluti…

Feature Importance

Missingness Bias in Model Debugging

2022-04-19 · ICLR 2022 4 · Saachi Jain, Hadi Salman, Eric Wong, Pengchuan Zhang 외

Missingness, or the absence of features from an input, is a concept fundamental to many model debugging tools. However, in computer vision, pixels cannot simply be removed from an image. One thus tends to resort to heuri…

model

To Impute or not to Impute? Missing Data in Treatment Effect Estimation

2022-02-04 · Jeroen Berrevoets, Fergus Imrie, Trent Kyono, James Jordon 외

Missing data is a systemic problem in practical scenarios that causes noise and bias when estimating treatment effects. This makes treatment effect estimation from data with missingness a particularly tricky endeavour. A…

Imputation

OverNaN: NaN-Aware Oversampling for Imbalanced Learning with Meaningful Missingness

2026-05-12 · Amanda S Barnard arxiv

Missing values are routinely treated as defects to be eliminated through deletion or imputation prior to machine learning. In many applied domains, however, missingness itself carries information, reflecting experimental…

Dimension Reduction for Data with Heterogeneous Missingness

2021-09-24 · Yurong Ling, Zijing Liu, Jing-Hao Xue

Dimension reduction plays a pivotal role in analysing high-dimensional data. However, observations with missing values present serious difficulties in directly applying standard dimension reduction techniques. As a large…

Dimensionality ReductionMissing Values