paper-with-me

Papers

High-Dimensional Linear Regression via Implicit Regularization

2019-03-22 · Peng Zhao, Yun Yang, Qiao-Chu He

Many statistical estimators for high-dimensional linear regression are M-estimators, formed through minimizing a data-dependent square loss function plus a regularizer. This work considers a new class of estimators implicitly defined through a discretized gradient dynamic system under overparameterization. We show that under suitable restricted isometry conditions, overparameterization leads to implicit regularization: if we directly apply gradient descent to the residual sum of squares with sufficiently small initial values, then under some proper early stopping rule, the iterates converge to a nearly sparse rate-optimal solution that improves over explicitly regularized approaches. In particular, the resulting estimator does not suffer from extra bias due to explicit penalties, and can achieve the parametric root-n rate when the signal-to-noise ratio is sufficiently high. We also perform simulations to compare our methods with high dimensional linear regression with explicit regularization. Our results illustrate the advantages of using implicit regularization via gradient descent after overparameterization in sparse vector estimation.

📄 PDF Abstract BibTeX arXiv:1903.09367

Code (0)

등록된 구현이 없습니다.

Tasks

regressionVocal Bursts Intensity Prediction

Methods 이 논문이 사용한 방법론

Early Stopping Early Stopping is a regularization technique for deep neural networks that stops training when parameter updates no longer begin to yield improves on a validation set. In…
Linear Regression Linear Regression is a method for modelling a relationship between a dependent variable and independent variables. These models can be fit with numerous approaches. The most…

Similar Papers 제목 키워드 기반

Optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization

2018-05-28 · Dmitry Kobak, Jonathan Lomond, Benoit Sanchez

A conventional wisdom in statistical learning is that large models require strong regularization to prevent overfitting. Here we show that this rule can be violated by linear regression in the underdetermined $n\ll p$ si…

Implicit Regularization for Group Sparsity

2023-01-29 · Jiangyuan Li, Thanh V. Nguyen, Chinmay Hegde, Raymond K. W. Wong

We study the implicit regularization of gradient descent towards structured sparsity via a novel neural reparameterization, which we call a diagonally grouped linear neural network. We show the following intriguing prope…

regression

Just Interpolate: Kernel "Ridgeless" Regression Can Generalize

2018-08-01 · Tengyuan Liang, Alexander Rakhlin

In the absence of explicit regularization, Kernel "Ridgeless" Regression with nonlinear kernels has the potential to fit the training data perfectly. It has been observed empirically, however, that such interpolated solu…

regression

The Benefits of Implicit Regularization from SGD in Least Squares Problems

2021-08-10 · NeurIPS 2021 12 · Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu 외

Stochastic gradient descent (SGD) exhibits strong algorithmic regularization effects in practice, which has been hypothesized to play an important role in the generalization of modern machine learning approaches. In this…

regression

Optimal L2 Regularization in High-dimensional Continual Linear Regression

2026-01-20 · Gilad Karpel, Edward Moroshko, Ran Levinstein, Ron Meir 외 arxiv

We study generalization in an overparameterized continual linear regression setting, where a model is trained with L2 (isotropic) regularization across a sequence of tasks. We derive a closed-form expression for the expe…

Continual Learning