paper-with-me

홈 › Papers

Toward Equation of Motion for Deep Neural Networks: Continuous-time Gradient Descent and Discretization Error Analysis

2022-10-28 · Taiki Miyagawa

We derive and solve an ``Equation of Motion'' (EoM) for deep neural networks (DNNs), a differential equation that precisely describes the discrete learning dynamics of DNNs. Differential equations are continuous but have played a prominent role even in the study of discrete optimization (gradient descent (GD) algorithms). However, there still exist gaps between differential equations and the actual learning dynamics of DNNs due to discretization error. In this paper, we start from gradient flow (GF) and derive a counter term that cancels the discretization error between GF and GD. As a result, we obtain EoM, a continuous differential equation that precisely describes the discrete learning dynamics of GD. We also derive discretization error to show to what extent EoM is precise. In addition, we apply EoM to two specific cases: scale- and translation-invariant layers. EoM highlights differences between continuous-time and discrete-time GD, indicating the importance of the counter term for a better description of the discrete learning dynamics of GD. Our experimental results support our theoretical findings.

📄 PDF Abstract BibTeX arXiv:2210.15898

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

EoM Excess of Mass aim to maximized the cluster stability

Similar Papers 제목 키워드 기반

A convergence analysis of the perturbed compositional gradient flow: averaging principle and normal deviations

2017-09-02 · Wenqing Hu, Chris Junchi Li

We consider in this work a system of two stochastic differential equations named the perturbed compositional gradient flow. By introducing a separation of fast and slow scales of the two equations, we show that the limit…

Towards Continuous-Time Approximations for Stochastic Gradient Descent without Replacement

2025-12-04 · Stefan Perko arxiv

Gradient optimization algorithms using epochs, that is those based on stochastic gradient descent without replacement (SGDo), are predominantly used to train machine learning models in practice. However, the mathematical…

Stochastic Gradient Descent in Continuous Time

2016-11-17 · Justin Sirignano, Konstantinos Spiliopoulos

Stochastic gradient descent in continuous time (SGDCT) provides a computationally efficient method for the statistical learning of continuous-time models, which are widely used in science, engineering, and finance. The S…

Effective continuous equations for adaptive SGD: a stochastic analysis view

2025-09-25 · Luca Callisti, Marco Romito, Francesco Triggiano arxiv

We present a theoretical analysis of some popular adaptive Stochastic Gradient Descent (SGD) methods in the small learning rate regime. Using the stochastic modified equations framework introduced by Li et al., we derive…

Stochastic Gradient Descent in Continuous Time: A Central Limit Theorem

2017-10-11 · Justin Sirignano, Konstantinos Spiliopoulos

Stochastic gradient descent in continuous time (SGDCT) provides a computationally efficient method for the statistical learning of continuous-time models, which are widely used in science, engineering, and finance. The S…