paper-with-me

Papers

De-biasing the Lasso: Optimal Sample Size for Gaussian Designs

2015-08-11 · Adel Javanmard, Andrea Montanari

Performing statistical inference in high-dimension is an outstanding challenge. A major source of difficulty is the absence of precise information on the distribution of high-dimensional estimators. Here, we consider linear regression in the high-dimensional regime $p\gg n$. In this context, we would like to perform inference on a high-dimensional parameters vector $\theta^*\in{\mathbb R}^p$. Important progress has been achieved in computing confidence intervals for single coordinates $\theta^*_i$. A key role in these new methods is played by a certain debiased estimator $\hat{\theta}^{\rm d}$ that is constructed from the Lasso. Earlier work establishes that, under suitable assumptions on the design matrix, the coordinates of $\hat{\theta}^{\rm d}$ are asymptotically Gaussian provided $\theta^*$ is $s_0$-sparse with $s_0 = o(\sqrt{n}/\log p )$. The condition $s_0 = o(\sqrt{n}/ \log p )$ is stronger than the one for consistent estimation, namely $s_0 = o(n/ \log p)$. We study Gaussian designs with known or unknown population covariance. When the covariance is known, we prove that the debiased estimator is asymptotically Gaussian under the nearly optimal condition $s_0 = o(n/ (\log p)^2)$. Note that earlier work was limited to $s_0 = o(\sqrt{n}/\log p)$ even for perfectly known covariance. The same conclusion holds if the population covariance is unknown but can be estimated sufficiently well, e.g. under the same sparsity conditions on the inverse covariance as assumed by earlier work. For intermediate regimes, we describe the trade-off between sparsity in the coefficients and in the inverse covariance of the design. We further discuss several applications of our results to high-dimensional inference. In particular, we propose a new estimator that is minimax optimal up to a factor $1+o_n(1)$ for i.i.d. Gaussian designs.

📄 PDF Abstract BibTeX arXiv:1508.02757

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Online Debiasing for Adaptively Collected High-dimensional Data with Applications to Time Series Analysis

2019-11-04 · Yash Deshpande, Adel Javanmard, Mohammad Mehrabi

Adaptive collection of data is commonplace in applications throughout science and engineering. From the point of view of statistical inference however, adaptive data collection induces memory and correlation in the sampl…

Time SeriesTime Series Analysis

Hypothesis Testing in High-Dimensional Regression under the Gaussian Random Design Model: Asymptotic Theory

2013-01-17 · Adel Javanmard, Andrea Montanari

We consider linear regression in the high-dimensional regime where the number of observations $n$ is smaller than the number of parameters $p$. A very successful approach in this setting uses $\ell_1$-penalized least squ…

Model SelectionregressionTwo-sample testing

Ridge Regression Revisited: Debiasing, Thresholding and Bootstrap

2020-09-17 · Yunyi Zhang, Dimitris N. Politis

The success of the Lasso in the era of high-dimensional data can be attributed to its conducting an implicit model selection, i.e., zeroing out regression coefficients that are not significant. By contrast, classical rid…

Model SelectionPrediction Intervalsregression

Fast Debiasing of the LASSO Estimator

2025-02-27 · Shuvayan Banerjee, James Saunderson, Radhendushka Srivastava, Ajit Rajwade

In high-dimensional sparse regression, the \textsc{Lasso} estimator offers excellent theoretical guarantees but is well-known to produce biased estimates. To address this, \cite{Javanmard2014} introduced a method to ``de…

L1-Regularized Least Squares for Support Recovery of High Dimensional Single Index Models with Gaussian Designs

2015-11-25 · Matey Neykov, Jun S. Liu, Tianxi Cai

It is known that for a certain class of single index models (SIMs) $Y = f(\boldsymbol{X}_{p \times 1}^\intercal\boldsymbol{\beta}_0, \varepsilon)$, support recovery is impossible when $\boldsymbol{X} \sim \mathcal{N}(0, …