paper-with-me

홈 › Papers

On the Power of Over-parametrization in Neural Networks with Quadratic Activation

2018-07-01 · ICML 2018 7 · Simon Du, Jason Lee

We provide new theoretical insights on why over-parametrization is effective in learning neural networks. For a $k$ hidden node shallow network with quadratic activation and $n$ training data points, we show as long as $ k \ge \sqrt{2n}$, over-parametrization enables local search algorithms to find a globally optimal solution for general smooth and convex loss functions. Further, despite that the number of parameters may exceed the sample size, using theory of Rademacher complexity, we show with weight decay, the solution also generalizes well if the data is sampled from a regular distribution such as Gaussian. To prove when $k\ge \sqrt{2n}$, the loss function has benign landscape properties, we adopt an idea from smoothed analysis, which may have other applications in studying loss surfaces of neural networks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Power of Over-parametrization in Neural Networks with Quadratic Activation

2018-03-03 · ICML 2018 · Simon S. Du, Jason D. Lee

We provide new theoretical insights on why over-parametrization is effective in learning neural networks. For a $k$ hidden node shallow network with quadratic activation and $n$ training data points, we show as long as $…

Understanding How Over-Parametrization Leads to Acceleration: A case of learning a single teacher neuron

2020-10-04 · Jun-Kun Wang, Jacob Abernethy

Over-parametrization has become a popular technique in deep learning. It is observed that by over-parametrization, a larger neural network needs a fewer training iterations than a smaller one to achieve a certain level o…

Quadratic Suffices for Over-parametrization via Matrix Chernoff Bound

2019-06-09 · Zhao Song, Xin Yang

We improve the over-parametrization size over two beautiful results [Li and Liang' 2018] and [Du, Zhai, Poczos and Singh' 2019] in deep learning theory.

Deep LearningLearning Theory

Nearly Minimal Over-Parametrization of Shallow Neural Networks

2019-10-09 · Armin Eftekhari, ChaeHwan Song, Volkan Cevher

A recent line of work has shown that an overparametrized neural network can perfectly fit the training data, an otherwise often intractable nonconvex optimization problem. For (fully-connected) shallow networks, in the b…

No bad local minima: Data independent training error guarantees for multilayer neural networks

2016-05-26 · Daniel Soudry, Yair Carmon

We use smoothed analysis techniques to provide guarantees on the training loss of Multilayer Neural Networks (MNNs) at differentiable local minima. Specifically, we examine MNNs with piecewise linear activation functions…