paper-with-me

홈 › Papers

Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK

2020-07-09 · Yuanzhi Li, Tengyu Ma, Hongyang R. Zhang

We consider the dynamic of gradient descent for learning a two-layer neural network. We assume the input $x\in\mathbb{R}^d$ is drawn from a Gaussian distribution and the label of $x$ satisfies $f^{\star}(x) = a^{\top}|W^{\star}x|$, where $a\in\mathbb{R}^d$ is a nonnegative vector and $W^{\star} \in\mathbb{R}^{d\times d}$ is an orthonormal matrix. We show that an over-parametrized two-layer neural network with ReLU activation, trained by gradient descent from random initialization, can provably learn the ground truth network with population loss at most $o(1/d)$ in polynomial time with polynomial samples. On the other hand, we prove that any kernel method, including Neural Tangent Kernel, with a polynomial number of samples in $d$, has population loss at least $\Omega(1 / d)$.

📄 PDF Abstract BibTeX arXiv:2007.04596

Code (0)

등록된 구현이 없습니다.

Tasks

Vocal Bursts Valence Prediction

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Nonasymptotic theory for two-layer neural networks: Beyond the bias-variance trade-off

2021-06-09 · Huiyuan Wang, Wei Lin

Large neural networks have proved remarkably effective in modern deep learning practice, even in the overparametrized regime where the number of active parameters is large relative to the sample size. This contradicts th…

Vocal Bursts Valence Prediction

Bayesian Free Energy of Deep ReLU Neural Network in Overparametrized Cases

2023-03-28 · Shuya Nagayasu, Sumio Watanabe

In many research fields in artificial intelligence, it has been shown that deep neural networks are useful to estimate unknown functions on high dimensional input spaces. However, their generalization performance is not …

Learning Theory

Memory capacity of neural networks with threshold and ReLU activations

2020-01-20 · Roman Vershynin

Overwhelming theoretical and empirical evidence shows that mildly overparametrized neural networks -- those with more connections than the size of the training data -- are often able to memorize the training data with $1…

Open-Ended Question Answering

Gradient representations in ReLU networks as similarity functions

2021-10-26 · Dániel Rácz, Bálint Daróczy

Feed-forward networks can be interpreted as mappings with linear decision surfaces at the level of the last layer. We investigate how the tangent space of the network can be exploited to refine the decision in case of Re…

Effect of Activation Functions on the Training of Overparametrized Neural Nets

2019-08-16 · ICLR 2020 1 · Abhishek Panigrahi, Abhishek Shetty, Navin Goyal

It is well-known that overparametrized neural networks trained using gradient-based methods quickly achieve small training error with appropriate hyperparameter settings. Recent papers have proved this statement theoreti…

Small Data Image Classification