A Double Robust Approach for Non-Monotone Missingness in Multi-Stage Data
Multivariate missingness with a non-monotone missing pattern is complicated to deal with in empirical studies. The traditional Missing at Random (MAR) assumption is difficult to justify in such cases. Previous studies have strengthened the MAR assumption, suggesting that the missing mechanism of any variable is random when conditioned on a uniform set of fully observed variables. However, empirical evidence indicates that this assumption may be violated for variables collected at different stages. This paper proposes a new MAR-type assumption that fits non-monotone missing scenarios involving multi-stage variables. Based on this assumption, we construct an Augmented Inverse Probability Weighted GMM (AIPW-GMM) estimator. This estimator features an asymmetric format for the augmentation term, guarantees double robustness, and achieves the closed-form semiparametric efficiency bound. We apply this method to cases of missingness in both endogenous regressor and outcome, using the Oregon Health Insurance Experiment as an example. We check the correlation between missing probabilities and partially observed variables to justify the assumption. Moreover, we find that excluding incomplete data results in a loss of efficiency and insignificant estimators. The proposed estimator reduces the standard error by more than 50% for the estimated effects of the Oregon Health Plan on the elderly.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Identification and Estimation of Long-Term Treatment Effects with Monotone Missing
Estimating long-term treatment effects has a wide range of applications in various domains. A key feature in this context is that collecting long-term outcomes typically involves a multi-stage process and is subject to m…
ImputationGenerative Modeling under Non-Monotone MAR Missingness via Approximate Wasserstein Gradient Flows
The prevalence of missing values in data science poses a substantial risk to any further analyses. Despite a wealth of research, principled nonparametric methods to deal with general non-monotone missingness are still sc…
Off-Policy Evaluation Under Nonignorable Missing Data
Off-Policy Evaluation (OPE) aims to estimate the value of a target policy using offline data collected from potentially different policies. In real-world applications, however, logged data often suffers from missingness.…
Imputation-Powered Inference
Modern multi-modal and multi-site data frequently suffer from blockwise missingness, where subsets of features are missing for groups of individuals, creating complex patterns that challenge standard inference methods. E…
Parallel Double Greedy Submodular Maximization
Many machine learning problems can be reduced to the maximization of submodular functions. Although well understood in the serial setting, the parallel maximization of submodular functions remains an open area of researc…