paper-with-me

홈 › Papers

Nearly Minimal Over-Parametrization of Shallow Neural Networks

2019-10-09 · Armin Eftekhari, ChaeHwan Song, Volkan Cevher

A recent line of work has shown that an overparametrized neural network can perfectly fit the training data, an otherwise often intractable nonconvex optimization problem. For (fully-connected) shallow networks, in the best case scenario, the existing theory requires quadratic over-parametrization as a function of the number of training samples. This paper establishes that linear overparametrization is sufficient to fit the training data, using a simple variant of the (stochastic) gradient descent. Crucially, unlike several related works, the training considered in this paper is not limited to the lazy regime in the sense cautioned against in [1, 2]. Beyond shallow networks, the framework developed in this work for over-parametrization is applicable to a variety of learning problems.

📄 PDF Abstract BibTeX arXiv:1910.03948

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Parametrization of subgrid scales in long-term simulations of the shallow-water equations using machine learning and convex limiting

2026-01-30 · Md Amran Hossan Mojamder, Zhihang Xu, Min Wang, Ilya Timofeyev arxiv

We present a method for parametrizing sub-grid processes in the Shallow Water equations. We define coarse variables and local spatial averages and use a feed-forward neural network to learn sub-grid fluxes. Our method re…

Shallow and Deep Networks are Near-Optimal Approximators of Korobov Functions

2021-09-29 · ICLR 2022 4 · Moise Blanchard, Mohammed Amine Bennouna

In this paper, we analyze the number of neurons and training parameters that a neural network needs to approximate multivariate functions of bounded second mixed derivatives --- Korobov functions. We prove upper bounds o…

Deep orthogonal linear networks are shallow

2020-11-27 · Pierre Ablin

We consider the problem of training a deep orthogonal linear network, which consists of a product of orthogonal matrices, with no non-linearity in-between. We show that training the weights with Riemannian gradient desce…

On the Power of Over-parametrization in Neural Networks with Quadratic Activation

2018-03-03 · ICML 2018 · Simon S. Du, Jason D. Lee

We provide new theoretical insights on why over-parametrization is effective in learning neural networks. For a $k$ hidden node shallow network with quadratic activation and $n$ training data points, we show as long as $…

On the Power of Over-parametrization in Neural Networks with Quadratic Activation

2018-07-01 · ICML 2018 7 · Simon Du, Jason Lee

We provide new theoretical insights on why over-parametrization is effective in learning neural networks. For a $k$ hidden node shallow network with quadratic activation and $n$ training data points, we show as long…