paper-with-me

Papers

High Dimensional M-Estimation with Missing Outcomes: A Semi-Parametric Framework

2019-11-26 · Abhishek Chakrabortty, Jiarui Lu, T. Tony Cai, Hongzhe Li

We consider high dimensional $M$-estimation in settings where the response $Y$ is possibly missing at random and the covariates $\mathbf{X} \in \mathbb{R}^p$ can be high dimensional compared to the sample size $n$. The parameter of interest $\boldsymbol{\theta}_0 \in \mathbb{R}^d$ is defined as the minimizer of the risk of a convex loss, under a fully non-parametric model, and $\boldsymbol{\theta}_0$ itself is high dimensional which is a key distinction from existing works. Standard high dimensional regression and series estimation with possibly misspecified models and missing $Y$ are included as special cases, as well as their counterparts in causal inference using 'potential outcomes'. Assuming $\boldsymbol{\theta}_0$ is $s$-sparse ($s \ll n$), we propose an $L_1$-regularized debiased and doubly robust (DDR) estimator of $\boldsymbol{\theta}_0$ based on a high dimensional adaptation of the traditional double robust (DR) estimator's construction. Under mild tail assumptions and arbitrarily chosen (working) models for the propensity score (PS) and the outcome regression (OR) estimators, satisfying only some high-level conditions, we establish finite sample performance bounds for the DDR estimator showing its (optimal) $L_2$ error rate to be $\sqrt{s (\log d)/ n}$ when both models are correct, and its consistency and DR properties when only one of them is correct. Further, when both the models are correct, we propose a desparsified version of our DDR estimator that satisfies an asymptotic linear expansion and facilitates inference on low dimensional components of $\boldsymbol{\theta}_0$. Finally, we discuss various of choices of high dimensional parametric/semi-parametric working models for the PS and OR estimators. All results are validated via detailed simulations.

📄 PDF Abstract BibTeX arXiv:1911.11345

Code (0)

등록된 구현이 없습니다.

Tasks

Causal InferenceregressionVocal Bursts Intensity Prediction

Methods 이 논문이 사용한 방법론

Causal inference Causal inference is the process of drawing a conclusion about a causal connection based on the conditions of the occurrence of an effect. The main difference between causal…

Similar Papers 제목 키워드 기반

PUATE: Efficient Average Treatment Effect Estimation from Treated (Positive) and Unlabeled Units

2025-01-31 · Masahiro Kato, Fumiaki Kozai, Ryo Inokuchi

The estimation of average treatment effects (ATEs), defined as the difference in expected outcomes between treatment and control groups, is a central topic in causal inference. This study develops semiparametric efficien…

Causal InferenceWeakly-supervised Learning

Deep Generalised Mixed Models: a Novel Neural Network Structure for Analysing Hierarchical Data

2026-08-06 · Nina van Gerwen, Dimitris Rizopoulos, Manon Hillegers, Loes Keijsers 외 arxiv

The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multiple times a day. Our work is motivated by such data collected by the G…

Data Augmentation

Likelihood Estimation with Incomplete Array Variate Observations

2012-09-12 · Deniz Akdemir

Missing data is an important challenge when dealing with high dimensional data arranged in the form of an array. In this paper, we propose methods for estimation of the parameters of array variate normal probability mode…

Imputation

Missing Data Estimation in High-Dimensional Datasets: A Swarm Intelligence-Deep Neural Network Approach

2016-07-01 · Collins Leke, Tshilidzi Marwala

In this paper, we examine the problem of missing data in high-dimensional datasets by taking into consideration the Missing Completely at Random and Missing at Random mechanisms, as well as theArbitrary missing pattern. …

Deep Learning

Rate Optimal Estimation and Confidence Intervals for High-dimensional Regression with Missing Covariates

2017-02-09 · Yining Wang, Jialei Wang, Sivaraman Balakrishnan, Aarti Singh

Although a majority of the theoretical literature in high-dimensional statistics has focused on settings which involve fully-observed data, settings with missing values and corruptions are common in practice. We consider…

Missing Valuesregression