paper-with-me

홈 › Papers

Gradient Descent on Neurons and its Link to Approximate Second-Order Optimization

2022-01-28 · Frederik Benzing

Second-order optimizers are thought to hold the potential to speed up neural network training, but due to the enormous size of the curvature matrix, they typically require approximations to be computationally tractable. The most successful family of approximations are Kronecker-Factored, block-diagonal curvature estimates (KFAC). Here, we combine tools from prior work to evaluate exact second-order updates with careful ablations to establish a surprising result: Due to its approximations, KFAC is not closely related to second-order updates, and in particular, it significantly outperforms true second-order updates. This challenges widely held believes and immediately raises the question why KFAC performs so well. Towards answering this question we present evidence strongly suggesting that KFAC approximates a first-order algorithm, which performs gradient descent on neurons rather than weights. Finally, we show that this optimizer often improves over KFAC in terms of computational cost and data-efficiency.

📄 PDF Abstract BibTeX arXiv:2201.12250

Code (1)

freedbee/neuron_descent_and_kfac 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Quadratic integrate-and-fire neurons exhibit less fragmented loss landscapes and outperform leaky integrate-and-fire neurons in spike-based gradient descent

2026-06-02 · Carlo Wenig, Raoul-Martin Memmesheimer, Christian Klos arxiv

The ability to train spiking neural networks is essential for modeling biological neural networks as well as for neuromorphic computing. However, for the extensively used leaky integrate-and-fire (LIF) neurons, arbitrari…

Splitting Steepest Descent for Growing Neural Architectures

2019-10-06 · NeurIPS 2019 12 · Qiang Liu, Lemeng Wu, Dilin Wang

We develop a progressive training approach for neural networks which adaptively grows the network structure by splitting existing neurons to multiple off-springs. By leveraging a functional steepest descent idea, we deri…

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent

2026-05-18 · Behrad Moniri, Hamed Hassani arxiv

We study feature learning in two-layer neural networks within the linear-width regime, where the number of hidden neurons, sample size, and input dimension scale proportionally. While recent work has analyzed feature lea…

The Quenching-Activation Behavior of the Gradient Descent Dynamics for Two-layer Neural Network Models

2020-06-25 · Chao Ma, Lei Wu, Weinan E

A numerical and phenomenological study of the gradient descent (GD) algorithm for training two-layer neural network models is carried out for different parameter regimes when the target function can be accurately approxi…

Long-time dynamics and universality of nonconvex gradient descent

2025-09-14 · Qiyang Han arxiv

This paper develops a general approach to characterize the long-time trajectory behavior of nonconvex gradient descent in generalized single-index models in the large aspect ratio regime. In this regime, we show that for…