paper-with-me

홈 › Papers

Towards Continuous-Time Approximations for Stochastic Gradient Descent without Replacement

2025-12-04 · Stefan Perko arxiv

Gradient optimization algorithms using epochs, that is those based on stochastic gradient descent without replacement (SGDo), are predominantly used to train machine learning models in practice. However, the mathematical theory of SGDo and related algorithms remain underexplored compared to their "with replacement" and "one-pass" counterparts. In this article, we propose a stochastic, continuous-time approximation to SGDo with additive noise based on a Young differential equation driven by a stochastic process we call an "epoched Brownian motion". We show its usefulness by proving the almost sure convergence of the continuous-time approximation for strongly convex objectives and learning rate schedules of the form $u_t = \frac{1}{(1+t)^β}, β\in (0,1)$. Moreover, we compute an upper bound on the asymptotic rate of almost sure convergence, which is as good or better than previous results for SGDo.

📄 PDF Abstract BibTeX arXiv:2512.04703

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Continuous Time Analysis of Momentum Methods

2019-06-10 · Nikola B. Kovachki, Andrew M. Stuart

Gradient descent-based optimization methods underpin the parameter training of neural networks, and hence comprise a significant component in the impressive test results found in a number of applications. Introducing sto…

First and Second Order Approximations to Stochastic Gradient Descent Methods with Momentum Terms

2025-04-18 · Eric Lu

Stochastic Gradient Descent (SGD) methods see many uses in optimization problems. Modifications to the algorithm, such as momentum-based SGD methods have been known to produce better results in certain cases. Much of thi…

Stochastic Modified Equations and Dynamics of Stochastic Gradient Algorithms I: Mathematical Foundations

2018-11-05 · Qianxiao Li, Cheng Tai, Weinan E

We develop the mathematical foundations of the stochastic modified equations (SME) framework for analyzing the dynamics of stochastic gradient algorithms, where the latter is approximated by a class of stochastic differe…

Emergence of heavy tails in homogenized stochastic gradient descent

2024-02-02 · Zhe Jiao, Martin Keller-Ressel

It has repeatedly been observed that loss minimization by stochastic gradient descent (SGD) leads to heavy-tailed distributions of neural network parameters. Here, we analyze a continuous diffusion approximation of SGD, …

An Inertial Newton Algorithm for Deep Learning

2019-05-29 · Camille Castera, Jérôme Bolte, Cédric Févotte, Edouard Pauwels

We introduce a new second-order inertial optimization method for machine learning called INNA. It exploits the geometry of the loss function while only requiring stochastic approximations of the function values and the g…

Deep LearningGeneral ClassificationImage Classification