paper-with-me

홈 › Papers

Singular-limit analysis of gradient descent with noise injection

2024-04-18 · Anna Shalova, André Schlichting, Mark Peletier

We study the limiting dynamics of a large class of noisy gradient descent systems in the overparameterized regime. In this regime the set of global minimizers of the loss is large, and when initialized in a neighbourhood of this zero-loss set a noisy gradient descent algorithm slowly evolves along this set. In some cases this slow evolution has been related to better generalisation properties. We characterize this evolution for the broad class of noisy gradient descent systems in the limit of small step size. Our results show that the structure of the noise affects not just the form of the limiting process, but also the time scale at which the evolution takes place. We apply the theory to Dropout, label noise and classical SGD (minibatching) noise, and show that these evolve on different two time scales. Classical SGD even yields a trivial evolution on both time scales, implying that additional noise is required for regularization. The results are inspired by the training of neural networks, but the theorems apply to noisy gradient descent of any loss that has a non-trivial zero-loss set.

📄 PDF Abstract BibTeX arXiv:2404.12293

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

An analysis on negative curvature induced by singularity in multi-layer neural-network learning

2010-12-01 · NeurIPS 2010 12 · Eiji Mizutani, Stuart Dreyfus

In the neural-network parameter space, an attractive field is likely to be induced by singularities. In such a singularity region, first-order gradient learning typically causes a long plateau with very little change i…

Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent

2024-02-04 · Naoki Sato, Hideaki Iiduka

For nonconvex objective functions, including those found in training deep neural networks, stochastic gradient descent (SGD) with momentum is said to converge faster and have better generalizability than SGD without mome…

Re-examining Low Rank adaptation for private LLM fine-tuning

2025-10-01 · Ali Dadsetan, Frank Rudzicz arxiv

Privacy is a central concern when fine-tuning large language models (LLMs) on sensitive data, and differentially private stochastic gradient descent (DP-SGD) -- which clips per-sample gradients and adds calibrated Gaussi…

Text Generation

Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization

2025-04-13 · NeurIPS 2021 12 · Gavin Zhang, Salar Fattahi, Richard Y. Zhang

In practical instances of nonconvex matrix factorization, the rank of the true solution $r^{\star}$ is often unknown, so the rank $r$ of the model can be overspecified as $r>r^{\star}$. This over-parameterized regime of …

De-singularity Subgradient for the $q$-th-Powered $\ell_p$-Norm Weber Location Problem

2024-12-20 · Zhao-Rong Lai, Xiaotian Wu, Liangda Fang, Ziliang Chen 외

The Weber location problem is widely used in several artificial intelligence scenarios. However, the gradient of the objective does not exist at a considerable set of singular points. Recently, a de-singularity subgradie…