paper-with-me

홈 › Papers

The instabilities of large learning rate training: a loss landscape view

2023-07-22 · Lawrence Wang, Stephen Roberts

Modern neural networks are undeniably successful. Numerous works study how the curvature of loss landscapes can affect the quality of solutions. In this work we study the loss landscape by considering the Hessian matrix during network training with large learning rates - an attractive regime that is (in)famously unstable. We characterise the instabilities of gradient descent, and we observe the striking phenomena of \textit{landscape flattening} and \textit{landscape shift}, both of which are intimately connected to the instabilities of training.

📄 PDF Abstract BibTeX arXiv:2307.11948

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Training Instabilities Induce Flatness Bias in Gradient Descent

2025-11-16 · Lawrence Wang, Stephen J. Roberts arxiv

Classical analyses of gradient descent (GD) define a stability threshold based on the largest eigenvalue of the loss Hessian, often termed sharpness. When the learning rate lies below this threshold, training is stable a…

Can Stability be Detrimental? Better Generalization through Gradient Descent Instabilities

2024-12-23 · Lawrence Wang, Stephen J. Roberts

Traditional analyses of gradient descent optimization show that, when the largest eigenvalue of the loss Hessian - often referred to as the sharpness - is below a critical learning-rate threshold, then training is 'stabl…

Tilting the playing field: Dynamical loss functions for machine learning

2021-02-07 · Miguel Ruiz-Garcia, Ge Zhang, Samuel S. Schoenholz, Andrea J. Liu

We show that learning can be improved by using loss functions that evolve cyclically during training to emphasize one class at a time. In underparameterized networks, such dynamical loss functions can lead to successful …

BIG-bench Machine Learning

Dynamical loss functions shape landscape topography and improve learning in artificial neural networks

2024-10-14 · Eduardo Lavin, Miguel Ruiz-Garcia

Dynamical loss functions are derived from standard loss functions used in supervised classification tasks, but they are modified such that the contribution from each class periodically increases and decreases. These osci…

Small-scale proxies for large-scale Transformer training instabilities

2023-09-25 · Mitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie Everett 외

Teams that have trained large Transformer-based models have reported training instabilities at large scale that did not appear when training with the same hyperparameters at smaller scales. Although the causes of such in…