paper-with-me

홈 › Papers

Early Stage Convergence and Global Convergence of Training Mildly Parameterized Neural Networks

2022-06-05 · Mingze Wang, Chao Ma

The convergence of GD and SGD when training mildly parameterized neural networks starting from random initialization is studied. For a broad range of models and loss functions, including the most commonly used square loss and cross entropy loss, we prove an `early stage convergence'' result. We show that the loss is decreased by a significant amount in the early stage of the training, and this decrease is fast. Furthurmore, for exponential type loss functions, and under some assumptions on the training data, we show global convergence of GD. Instead of relying on extreme over-parameterization, our study is based on a microscopic analysis of the activation patterns for the neurons, which helps us derive more powerful lower bounds for the gradient. The results on activation patterns, which we call `neuron partition'', help build intuitions for understanding the behavior of neural networks' training dynamics, and may be of independent interest.

📄 PDF Abstract BibTeX arXiv:2206.02139

Code (1)

wmz9/early_stage_convergence_neurips2022 공식 구현

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Accelerating Frank-Wolfe via Averaging Step Directions

2022-05-24 · Zhaoyue Chen, Yifan Sun

The Frank-Wolfe method is a popular method in sparse constrained optimization, due to its fast per-iteration complexity. However, the tradeoff is that its worst case global convergence is comparatively slow, and importan…

Global Update Guided Federated Learning

2022-04-08 · Qilong Wu, Lin Liu, Shibei Xue

Federated learning protects data privacy and security by exchanging models instead of data. However, unbalanced data distributions among participating clients compromise the accuracy and convergence speed of federated le…

Federated Learning

Towards Understanding Label Smoothing

2020-06-20 · Yi Xu, Yuanhong Xu, Qi Qian, Hao Li 외

Label smoothing regularization (LSR) has a great success in training deep neural networks by stochastic algorithms such as stochastic gradient descent and its variants. However, the theoretical understanding of its power…

Multi-stage, multi-swarm PSO for joint optimization of well placement and control

2021-06-02 · Ajitabh Kumar

Evolutionary optimization algorithms, including particle swarm optimization (PSO), have been successfully applied in oil industry for production planning and control. Such optimization studies are quite challenging due t…

Global Convergence and Geometric Characterization of Slow to Fast Weight Evolution in Neural Network Training for Classifying Linearly Non-Separable Data

2020-02-28 · Ziang Long, Penghang Yin, Jack Xin

In this paper, we study the dynamics of gradient descent in learning neural networks for classification problems. Unlike in existing works, we consider the linearly non-separable case where the training data of different…

General Classification