paper-with-me

Papers

On the Convergence of Gradient Descent Training for Two-layer ReLU-networks in the Mean Field Regime

2020-05-27 · Stephan Wojtowytsch

We describe a necessary and sufficient condition for the convergence to minimum Bayes risk when training two-layer ReLU-networks by gradient descent in the mean field regime with omni-directional initial parameter distribution. This article extends recent results of Chizat and Bach to ReLU-activated networks and to the situation in which there are no parameters which exactly achieve MBR. The condition does not depend on the initalization of parameters and concerns only the weak convergence of the realization of the neural network, not its parameter distribution.

📄 PDF Abstract BibTeX arXiv:2005.13530

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Proof of Global Convergence of Gradient Descent for Deep ReLU Networks with Linear Widths

2021-01-24 · Quynh Nguyen

We give a simple proof for the global convergence of gradient descent in training deep ReLU networks with the standard square loss, and show some of its improvements over the state-of-the-art. In particular, while prior …

Magnitude and Angle Dynamics in Training Single ReLU Neurons

2022-09-27 · Sangmin Lee, Byeongsu Sim, Jong Chul Ye

To understand learning the dynamics of deep ReLU networks, we investigate the dynamic system of gradient flow $w(t)$ by decomposing it to magnitude $w(t)$ and angle $\phi(t):= \pi - \theta(t) $ components. In particular,…

A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions

2021-04-01 · Arnulf Jentzen, Adrian Riekert

In this article we study the stochastic gradient descent (SGD) optimization method in the training of fully-connected feedforward artificial neural networks with ReLU activation. The main result of this work proves that …

Directional Convergence, Benign Overfitting of Gradient Descent in leaky ReLU two-layer Neural Networks

2025-05-22 · Ichiro Hashimoto

In this paper, we prove directional convergence of network parameters of fixed width leaky ReLU two-layer neural networks optimized by gradient descent with exponential loss, which was previously only known for gradient …

Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal Data

2023-10-29 · NeurIPS 2023 11

The implicit bias towards solutions with favorable properties is believed to be a key reason why neural networks trained by gradient-based optimization can generalize well. While the implicit bias of gradient flow has be…