paper-with-me

홈 › Papers

Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation

2025-02-10 · Matthew Esmaili Mallory, Kevin Han Huang, Morgane Austern

Over the last decade, a wave of research has characterized the exact asymptotic risk of many high-dimensional models in the proportional regime. Two foundational results have driven this progress: Gaussian universality, which shows that the asymptotic risk of estimators trained on non-Gaussian and Gaussian data is equivalent, and the convex Gaussian min-max theorem (CGMT), which characterizes the risk under Gaussian settings. However, these results rely on the assumption that the data consists of independent random vectors--an assumption that significantly limits its applicability to many practical setups. In this paper, we address this limitation by generalizing both results to the dependent setting. More precisely, we prove that Gaussian universality still holds for high-dimensional logistic regression under block dependence, $m$-dependence and special cases of mixing, and establish a novel CGMT framework that accommodates for correlation across both the covariates and observations. Using these results, we establish the impact of data augmentation, a widespread practice in deep learning, on the asymptotic risk.

📄 PDF Abstract BibTeX arXiv:2502.15752

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Methods 이 논문이 사용한 방법론

Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…

Similar Papers 제목 키워드 기반

High-dimensional logistic regression with missing data: Imputation, regularization, and universality

2024-10-01 · Kabir Aladin Verchand, Andrea Montanari

We study high-dimensional, ridge-regularized logistic regression in a setting in which the covariates may be missing or corrupted by additive noise. When both the covariates and the additive corruptions are independent a…

ImputationPredictionregression

A Model of Double Descent for High-dimensional Binary Linear Classification

2019-11-13 · Zeyu Deng, Abla Kammoun, Christos Thrampoulidis

We consider a model for logistic regression where only a subset of features of size $p$ is used for training a linear classifier over $n$ training samples. The classifier is obtained by running gradient descent (GD) on l…

ClassificationGeneral ClassificationregressionVocal Bursts Intensity Prediction

A Novel Gaussian Min-Max Theorem and its Applications

2024-02-12 · Danil Akhtiamov, David Bosch, Reza Ghane, K Nithin Varma 외

A celebrated result by Gordon allows one to compare the min-max behavior of two Gaussian processes if certain inequality conditions are met. The consequences of this result include the Gaussian min-max (GMT) and convex G…

Binary ClassificationGaussian Processes

RAPTOR: Ridge-Adaptive Logistic Probes

2026-01-29 · Ziqi Gao, Yaotian Zhu, Qingcheng Zeng, Xu Zhao 외 arxiv

Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used operationally in probe-then-steer pipelines: …

Characterization of Gaussian Universality Breakdown in High-Dimensional Empirical Risk Minimization

2026-04-03 · Chiheb Yaakoubi, Cosme Louart, Malik Tiomoko, Zhenyu Liao arxiv

We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs. By heuristically extending the Convex Gaussian Min-Max Theorem (CGMT) to non-Gaussian settings, we derive an asy…