paper-with-me

홈 › Papers

Global Convergence of Three-layer Neural Networks in the Mean Field Regime

2021-05-11 · ICLR 2021 1 · Huy Tuan Pham, Phan-Minh Nguyen

In the mean field regime, neural networks are appropriately scaled so that as the width tends to infinity, the learning dynamics tends to a nonlinear and nontrivial dynamical limit, known as the mean field limit. This lends a way to study large-width neural networks via analyzing the mean field limit. Recent works have successfully applied such analysis to two-layer networks and provided global convergence guarantees. The extension to multilayer ones however has been a highly challenging puzzle, and little is known about the optimization efficiency in the mean field regime when there are more than two layers. In this work, we prove a global convergence result for unregularized feedforward three-layer networks in the mean field regime. We first develop a rigorous framework to establish the mean field limit of three-layer networks under stochastic gradient descent training. To that end, we propose the idea of a \textit{neuronal embedding}, which comprises of a fixed probability space that encapsulates neural networks of arbitrary sizes. The identified mean field limit is then used to prove a global convergence guarantee under suitable regularity and convergence mode assumptions, which -- unlike previous works on two-layer networks -- does not rely critically on convexity. Underlying the result is a universal approximation property, natural of neural networks, which importantly is shown to hold at \textit{any} finite training time (not necessarily at convergence) via an algebraic topology argument.

📄 PDF Abstract BibTeX arXiv:2105.05228

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Note on the Global Convergence of Multilayer Neural Networks in the Mean Field Regime

2020-06-16 · Huy Tuan Pham, Phan-Minh Nguyen

In a recent work, we introduced a rigorous framework to describe the mean field limit of the gradient-based learning dynamics of multilayer neural networks, based on the idea of a neuronal embedding. There we also proved…

Diversity

Mean-Field Analysis of Two-Layer Neural Networks: Global Optimality with Linear Convergence Rates

2022-05-19 · Jingwei Zhang, Xunpeng Huang

We consider optimizing two-layer neural networks in the mean-field regime where the learning dynamics of network weights can be approximated by the evolution in the space of probability measures over the weight parameter…

Mean-field analysis for heavy ball methods: Dropout-stability, connectivity, and global convergence

2022-10-13 · Diyuan Wu, Vyacheslav Kungurtsev, Marco Mondelli

The stochastic heavy ball method (SHB), also known as stochastic gradient descent (SGD) with Polyak's momentum, is widely used in training neural networks. However, despite the remarkable success of such algorithm in pra…

A Rigorous Framework for the Mean Field Limit of Multilayer Neural Networks

2020-01-30 · Phan-Minh Nguyen, Huy Tuan Pham

We develop a mathematically rigorous framework for multilayer neural networks in the mean field regime. As the network's widths increase, the network's learning trajectory is shown to be well captured by a meaningful and…

Global Convergence of Second-order Dynamics in Two-layer Neural Networks

2020-07-14 · Walid Krichene, Kenneth F. Caluya, Abhishek Halder

Recent results have shown that for two-layer fully connected neural networks, gradient flow converges to a global optimum in the infinite width limit, by making a connection between the mean field dynamics and the Wasser…

Vocal Bursts Valence Prediction