paper-with-me

홈 › Papers

Bounds on Over-Parameterization for Guaranteed Existence of Descent Paths in Shallow ReLU Networks

2020-05-01 · ICLR 2020 1 · Arsalan Sharifnassab, Saber Salehkaleybar, S. Jamaloddin Golestani

We study the landscape of squared loss in neural networks with one-hidden layer and ReLU activation functions. Let $m$ and $d$ be the widths of hidden and input layers, respectively. We show that there exist poor local minima with positive curvature for some training sets of size $n\geq m+2d-2$. By positive curvature of a local minimum, we mean that within a small neighborhood the loss function is strictly increasing in all directions. Consequently, for such training sets, there are initialization of weights from which there is no descent path to global optima. It is known that for $n\le m$, there always exist descent paths to global optima from all initial weights. In this perspective, our results provide a somewhat sharp characterization of the over-parameterization required for "existence of descent paths" in the loss landscape.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…

Similar Papers 제목 키워드 기반

Over-parameterization Improves Generalization in the XOR Detection Problem

2019-05-01 · ICLR 2019 5 · Alon Brutzkus, Amir Globerson

Empirical evidence suggests that neural networks with ReLU activations generalize better with over-parameterization. However, there is currently no theoretical analysis that explains this observation. In this work, we st…

Direct Parameterization of Lipschitz-Bounded Deep Networks

2023-01-27 · Ruigang Wang, Ian R. Manchester

This paper introduces a new parameterization of deep neural networks (both fully-connected and convolutional) with guaranteed $\ell^2$ Lipschitz bounds, i.e. limited sensitivity to input perturbations. The Lipschitz guar…

image-classificationImage Classification

Over-Parameterization Exponentially Slows Down Gradient Descent for Learning a Single Neuron

2023-02-20 · Weihang Xu, Simon S. Du

We revisit the problem of learning a single neuron with ReLU activation under Gaussian input with square loss. We particularly focus on the over-parameterization setting where the student network has $n\ge 2$ neurons. We…

Beating the Perils of Non-Convexity: Guaranteed Training of Neural Networks using Tensor Methods

2015-06-28 · Majid Janzamin, Hanie Sedghi, Anima Anandkumar

Training neural networks is a challenging non-convex optimization problem, and backpropagation or gradient descent can get stuck in spurious local optima. We propose a novel algorithm based on tensor decomposition for gu…

Tensor Decomposition

Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks

2019-02-04 · Yuan Cao, Quanquan Gu

Empirical studies show that gradient-based methods can learn deep neural networks (DNNs) with very good generalization performance in the over-parameterization regime, where DNNs can easily fit a random labeling of the t…

Generalization Bounds