paper-with-me

홈 › Papers

The Quenching-Activation Behavior of the Gradient Descent Dynamics for Two-layer Neural Network Models

2020-06-25 · Chao Ma, Lei Wu, Weinan E

A numerical and phenomenological study of the gradient descent (GD) algorithm for training two-layer neural network models is carried out for different parameter regimes when the target function can be accurately approximated by a relatively small number of neurons. It is found that for Xavier-like initialization, there are two distinctive phases in the dynamic behavior of GD in the under-parametrized regime: An early phase in which the GD dynamics follows closely that of the corresponding random feature model and the neurons are effectively quenched, followed by a late phase in which the neurons are divided into two groups: a group of a few "activated" neurons that dominate the dynamics and a group of background (or "quenched") neurons that support the continued activation and deactivation process. This neural network-like behavior is continued into the mildly over-parametrized regime, where it undergoes a transition to a random feature-like behavior. The quenching-activation process seems to provide a clear mechanism for "implicit regularization". This is qualitatively different from the dynamics associated with the "mean-field" scaling where all neurons participate equally and there does not appear to be qualitative changes when the network parameters are changed.

📄 PDF Abstract BibTeX arXiv:2006.14450

Code (1)

TheoreticalML/GD.quenching_activation 공식 구현 pytorch

Similar Papers 제목 키워드 기반

On the asymptotics of wide networks with polynomial activations

2020-06-11 · Kyle Aitken, Guy Gur-Ari

We consider an existing conjecture addressing the asymptotic behavior of neural networks in the large width limit. The results that follow from this conjecture include tight bounds on the behavior of wide networks during…

The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System

2026-07-06 · Thomas Hofmann arxiv

Many phenomena of deep learning are dynamical: they concern not only which minima exist, but how gradient descent reaches, avoids, or selects among them. Edge-of-stability behavior, sharpness oscillations, catapult phase…

Towards an Understanding of Residual Networks Using Neural Tangent Hierarchy (NTH)

2020-07-07 · Yuqing Li, Tao Luo, Nung Kwan Yip

Gradient descent yields zero training loss in polynomial time for deep neural networks despite non-convex nature of the objective function. The behavior of network in the infinite width limit trained by gradient descent …

Convergence Results for Neural Networks via Electrodynamics

2017-02-01 · Rina Panigrahy, Sushant Sachdeva, Qiuyi Zhang

We study whether a depth two neural network can learn another depth two network using gradient descent. Assuming a linear output node, we show that the question of whether gradient descent converges to the target functio…

On the global convergence of gradient descent for wide shallow models with bounded nonlinearities

2026-05-11 · Romain Petit, Clarice Poon, Gabriel Peyré arxiv

A surprising phenomenon in the training of neural networks is the ability of gradient descent to find global minimizers of the training loss despite its non-convexity. Following earlier works, we investigate this behavio…