paper-with-me

홈 › Papers

Moving Beyond Sub-Gaussianity in High-Dimensional Statistics: Applications in Covariance Estimation and Linear Regression

2018-04-08 · Arun Kumar Kuchibhotla, Abhishek Chakrabortty

Concentration inequalities form an essential toolkit in the study of high dimensional (HD) statistical methods. Most of the relevant statistics literature in this regard is based on sub-Gaussian or sub-exponential tail assumptions. In this paper, we first bring together various probabilistic inequalities for sums of independent random variables under much more general exponential type (namely sub-Weibull) tail assumptions. These results extract a part sub-Gaussian tail behavior in finite samples, matching the asymptotics governed by the central limit theorem, and are compactly represented in terms of a new Orlicz quasi-norm - the Generalized Bernstein-Orlicz norm - that typifies such tail behaviors. We illustrate the usefulness of these inequalities through the analysis of four fundamental problems in HD statistics. In the first two problems, we study the rate of convergence of the sample covariance matrix in terms of the maximum elementwise norm and the maximum k-sub-matrix operator norm which are key quantities of interest in bootstrap, HD covariance matrix estimation and HD inference. The third example concerns the restricted eigenvalue condition, required in HD linear regression, which we verify for all sub-Weibull random vectors through a unified analysis, and also prove a more general result related to restricted strong convexity in the process. In the final example, we consider the Lasso estimator for linear regression and establish its rate of convergence under much weaker than usual tail assumptions (on the errors as well as the covariates), while also allowing for misspecified models and both fixed and random design. To our knowledge, these are the first such results for Lasso obtained in this generality. The common feature in all our results over all the examples is that the convergence rates under most exponential tails match the usual ones under sub-Gaussian assumptions.

📄 PDF Abstract BibTeX arXiv:1804.02605

Code (0)

등록된 구현이 없습니다.

Tasks

regression

Methods 이 논문이 사용한 방법론

Linear Regression Linear Regression is a method for modelling a relationship between a dependent variable and independent variables. These models can be fit with numerous approaches. The most…

Similar Papers 제목 키워드 기반

Approximate Gaussianity Beyond Initialisation in Neural Networks

2025-10-06 · Edward Hirst, Sanjaye Ramgoolam arxiv

Ensembles of neural network weight matrices are studied through the training process for the MNIST classification problem, testing the efficacy of matrix models for representing their distributions, under assumptions of …

Blind Determination of the Number of Sources Using Distance Correlation

2020-08-30 · Amir Weiss, Arie Yeredor

A novel blind estimate of the number of sources from noisy, linear mixtures is proposed. Based on Sz\'ekely et al.'s distance correlation measure, we define the Sources' Dependency Criterion (SDC), from which our estimat…

Moment- and Power-Spectrum-Based Gaussianity Regularization for Text-to-Image Models

2025-09-07 · Jisung Hwang, Jaihoon Kim, Minhyuk Sung arxiv

We propose a novel regularization loss that enforces standard Gaussianity, encouraging samples to align with a standard Gaussian distribution. This facilitates a range of downstream tasks involving optimization in the la…

Vector-valued self-normalized concentration inequalities beyond sub-Gaussianity

2025-11-05 · Diego Martinez-Taboada, Tomas Gonzalez, Aaditya Ramdas arxiv

The study of self-normalized processes plays a crucial role in a wide range of applications, from sequential decision-making to econometrics. While the behavior of self-normalized concentration has been widely investigat…

On the Identifiability of Sparse ICA without Assuming Non-Gaussianity

2024-08-19 · NeurIPS 2023 11 · Ignavier Ng, Yujia Zheng, Xinshuai Dong, Kun Zhang

Independent component analysis (ICA) is a fundamental statistical tool used to reveal hidden generative processes from observed data. However, traditional ICA approaches struggle with the rotational invariance inherent i…