paper-with-me

홈 › Papers

Gradient Descent as a Perceptron Algorithm: Understanding Dynamics and Implicit Acceleration

2025-12-12 · Alexander Tyurin arxiv

Even for the gradient descent (GD) method applied to neural network training, understanding its optimization dynamics, including convergence rate, iterate trajectories, function value oscillations, and especially its implicit acceleration, remains a challenging problem. We analyze nonlinear models with the logistic loss and show that the steps of GD reduce to those of generalized perceptron algorithms (Rosenblatt, 1958), providing a new perspective on the dynamics. This reduction yields significantly simpler algorithmic steps, which we analyze using classical linear algebra tools. Using these tools, we demonstrate on a minimalistic example that the nonlinearity in a two-layer model can provably yield a faster iteration complexity $\tilde{O}(\sqrt{d})$ compared to $Ω(d)$ achieved by linear models, where $d$ is the number of features. This helps explain the optimization dynamics and the implicit acceleration phenomenon observed in neural networks. The theoretical results are supported by extensive numerical experiments. We believe that this alternative view will further advance research on the optimization of neural networks.

📄 PDF Abstract BibTeX arXiv:2512.11587

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Blind Descent: A Prequel to Gradient Descent

2020-06-20 · Akshat Gupta, Prasad N R

We describe an alternative learning method for neural networks, which we call Blind Descent. By design, Blind Descent does not face problems like exploding or vanishing gradients. In Blind Descent, gradients are not used…

Quaternion MLP Neural Networks Based on the Maximum Correntropy Criterion

2023-09-11 · Gang Wang, Xinyu Tian, Zuxuan Zhang

We propose a gradient ascent algorithm for quaternion multilayer perceptron (MLP) networks based on the cost function of the maximum correntropy criterion (MCC). In the algorithm, we use the split quaternion activation f…

Surrogate Functions for Maximizing Precision at the Top

2015-05-26 · Purushottam Kar, Harikrishna Narasimhan, Prateek Jain

The problem of maximizing precision at the top of a ranked list, often dubbed Precision@k (prec@k), finds relevance in myriad learning applications such as ranking, multi-label classification, and learning with severe la…

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATION

On a continuous time model of gradient descent dynamics and instability in deep learning

2023-02-03 · Mihaela Rosca, Yan Wu, Chongli Qin, Benoit Dherin

The recipe behind the success of deep learning has been the combination of neural networks and gradient-based optimization. Understanding the behavior of gradient descent however, and particularly its instability, has la…

Deep Learning

Dynamic Decoupling of Placid Terminal Attractor-based Gradient Descent Algorithm

2024-09-10 · Jinwei Zhao, Marco Gori, Alessandro Betti, Stefano Melacci 외

Gradient descent (GD) and stochastic gradient descent (SGD) have been widely used in a large number of application domains. Therefore, understanding the dynamics of GD and improving its convergence speed is still of grea…

image-classificationImage Classification