paper-with-me

Papers

High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation

2022-05-03 · Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, Greg Yang

We study the first gradient descent step on the first-layer parameters $\boldsymbol{W}$ in a two-layer neural network: $f(\boldsymbol{x}) = \frac{1}{\sqrt{N}}\boldsymbol{a}^\top\sigma(\boldsymbol{W}^\top\boldsymbol{x})$, where $\boldsymbol{W}\in\mathbb{R}^{d\times N}, \boldsymbol{a}\in\mathbb{R}^{N}$ are randomly initialized, and the training objective is the empirical MSE loss: $\frac{1}{n}\sum_{i=1}^n (f(\boldsymbol{x}_i)-y_i)^2$. In the proportional asymptotic limit where $n,d,N\to\infty$ at the same rate, and an idealized student-teacher setting, we show that the first gradient update contains a rank-1 "spike", which results in an alignment between the first-layer weights and the linear component of the teacher model $f^*$. To characterize the impact of this alignment, we compute the prediction risk of ridge regression on the conjugate kernel after one gradient step on $\boldsymbol{W}$ with learning rate $\eta$, when $f^*$ is a single-index model. We consider two scalings of the first step learning rate $\eta$. For small $\eta$, we establish a Gaussian equivalence property for the trained feature map, and prove that the learned kernel improves upon the initial random features model, but cannot defeat the best linear model on the input. Whereas for sufficiently large $\eta$, we prove that for certain $f^*$, the same ridge estimator on trained features can go beyond this "linear regime" and outperform a wide range of random features and rotationally invariant kernels. Our results demonstrate that even one gradient step can lead to a considerable advantage over random features, and highlight the role of learning rate scaling in the initial phase of training.

📄 PDF Abstract BibTeX arXiv:2205.01445

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Asymptotics of feature learning in two-layer networks after one gradient-step

2024-02-07 · Hugo Cui, Luca Pesce, Yatin Dandi, Florent Krzakala 외

In this manuscript, we investigate the problem of how two-layer neural networks learn features from data, and improve over the kernel regime, after being trained with a single gradient descent step. Leveraging the insigh…

Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models

2025-02-07 · Alexander Atanasov, Blake Bordelon, Jacob A. Zavatone-Veth, Courtney Paquette 외

We derive a novel deterministic equivalence for the two-point function of a random matrix resolvent. Using this result, we give a unified derivation of the performance of a wide variety of high-dimensional linear models …

regression

Tuning Stochastic Gradient Algorithms for Statistical Inference via Large-Sample Asymptotics

2022-07-25 · Jeffrey Negrea, Jun Yang, Haoyue Feng, Daniel M. Roy 외

The tuning of stochastic gradient algorithms (SGAs) for optimization and sampling is often based on heuristics and trial-and-error rather than generalizable theory. We address this theory--practice gap by characterizing …

Nadaraya-Watson kernel smoothing as a random energy model

2024-08-07 · Jacob A. Zavatone-Veth, Cengiz Pehlevan

Precise asymptotics have revealed many surprises in high-dimensional regression. These advances, however, have not extended to perhaps the simplest estimator: direct Nadaraya-Watson (NW) kernel smoothing. Here, we descri…

SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality

2021-02-08 · Courtney Paquette, Kiwon Lee, Fabian Pedregosa, Elliot Paquette

We propose a new framework, inspired by random matrix theory, for analyzing the dynamics of stochastic gradient descent (SGD) when both number of samples and dimensions are large. This framework applies to any fixed step…