paper-with-me

홈 › Papers

Why Does Stochastic Gradient Descent Slow Down in Low-Precision Training?

2025-08-10 · Vincent-Daniel Yun arxiv

Low-precision training has become crucial for reducing the computational and memory costs of large-scale deep learning. However, quantizing gradients introduces magnitude shrinkage, which can change how stochastic gradient descent (SGD) converges. In this study, we explore SGD convergence under a gradient shrinkage model, where each stochastic gradient is scaled by a factor \( q_k \in (0,1] \). We show that this shrinkage affect the usual stepsize \( μ_k \) with an effective stepsize \( μ_k q_k \), slowing convergence when \( q_{\min} < 1 \). With typical smoothness and bounded-variance assumptions, we prove that low-precision SGD still converges, but at a slower pace set by \( q_{\min} \), and with a higher steady error level due to quantization effects. We analyze theoretically how lower numerical precision slows training by treating it as gradient shrinkage within the standard SGD convergence setup.

📄 PDF Abstract BibTeX arXiv:2508.07142

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Accelerating Stochastic Gradient Descent using Predictive Variance Reduction

2013-12-01 · NeurIPS 2013 12 · Rie Johnson, Tong Zhang

Stochastic gradient descent is popular for large scale optimization but has slow convergence asymptotically due to the inherent variance. To remedy this problem, we introduce an explicit variance reduction method for sto…

Structured Prediction

Frank-Wolfe optimization for deep networks

2020-06-06 · Jakob Stigenberg

Deep neural networks is today one of the most popular choices in classification, regression and function approximation. However, the training of such deep networks is far from trivial as there are often millions of param…

regression

A qualitative difference between gradient flows of convex functions in finite- and infinite-dimensional Hilbert spaces

2023-10-26 · Jonathan W. Siegel, Stephan Wojtowytsch

We consider gradient flow/gradient descent and heavy ball/accelerated gradient descent optimization for convex objective functions. In the gradient flow case, we prove the following: 1. If $f$ does not have a minimizer, …

Stochastic Gradient Descent: Going As Fast As Possible But Not Faster

2017-09-05 · Alice Schoenauer-Sebag, Marc Schoenauer, Michèle Sebag

When applied to training deep neural networks, stochastic gradient descent (SGD) often incurs steady progression phases, interrupted by catastrophic episodes in which loss and gradient norm explode. A possible mitigation…

Change Point Detection

Accelerating Stochastic Gradient Descent Using Antithetic Sampling

2018-10-07 · Jingchang Liu, Linli Xu

(Mini-batch) Stochastic Gradient Descent is a popular optimization method which has been applied to many machine learning applications. But a rather high variance introduced by the stochastic gradient in each step may sl…

Binary ClassificationGeneral Classification