paper-with-me

홈 › Papers

Understanding the Generalization Benefits of Late Learning Rate Decay

2024-01-21 · Yinuo Ren, Chao Ma, Lexing Ying

Why do neural networks trained with large learning rates for a longer time often lead to better generalization? In this paper, we delve into this question by examining the relation between training and testing loss in neural networks. Through visualization of these losses, we note that the training trajectory with a large learning rate navigates through the minima manifold of the training loss, finally nearing the neighborhood of the testing loss minimum. Motivated by these findings, we introduce a nonlinear model whose loss landscapes mirror those observed for real neural networks. Upon investigating the training process using SGD on our model, we demonstrate that an extended phase with a large learning rate steers our model towards the minimum norm solution of the training loss, which may achieve near-optimal generalization, thereby affirming the empirically observed benefits of late learning rate decay.

📄 PDF Abstract BibTeX arXiv:2401.11600

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Investigating the Role of Weight Decay in Enhancing Nonconvex SGD

2025-01-01 · CVPR 2025 1 · Tao Sun, Yuhao Huang, Li Shen, Kele Xu 외

Weight decay is a widely used technique in training machine learning models, known to empirically enhance the generalization of Stochastic Gradient Descent (SGD). While intuitively weight decay allows SGD to train a …

Gaussian Process Inference Using Mini-batch Stochastic Gradient Descent: Convergence Guarantees and Empirical Benefits

2021-11-19 · Hao Chen, Lili Zheng, Raed Al Kontar, Garvesh Raskutti

Stochastic gradient descent (SGD) and its variants have established themselves as the go-to algorithms for large-scale machine learning problems with independent samples due to their generalization performance and intrin…

Noisy Information Bottlenecks for Generalization

2018-09-27 · Julius Kunze, Louis Kirsch, Hippolyt Ritter, David Barber

We propose Noisy Information Bottlenecks (NIB) to limit mutual information between learned parameters and the data through noise. We show why this benefits generalization and allows mitigation of model overfitting both f…

Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks

2025-10-13 · Xuan Tang, Han Zhang, Yuan Cao, Difan Zou arxiv

Adam is a popular and widely used adaptive gradient method in deep learning, which has also received tremendous focus in theoretical research. However, most existing theoretical work primarily analyzes its full-batch ver…

Probability-Dependent Gradient Decay in Large Margin Softmax

2022-10-31 · Siyuan Zhang, Linbo Xie, Ying Chen

In the past few years, Softmax has become a common component in neural network frameworks. In this paper, a gradient decay hyperparameter is introduced in Softmax to control the probability-dependent gradient decay rate …