paper-with-me

홈 › Papers

Exploring loss function topology with cyclical learning rates

2017-02-14 · Leslie N. Smith, Nicholay Topin

We present observations and discussion of previously unreported phenomena discovered while training residual networks. The goal of this work is to better understand the nature of neural networks through the examination of these new empirical results. These behaviors were identified through the application of Cyclical Learning Rates (CLR) and linear network interpolation. Among these behaviors are counterintuitive increases and decreases in training loss and instances of rapid training. For example, we demonstrate how CLR can produce greater testing accuracy than traditional training despite using large learning rates. Files to replicate these results are available at https://github.com/lnsmith54/exploring-loss

📄 PDF Abstract BibTeX arXiv:1702.04283

Code (2)

lnsmith54/exploring-loss 공식 구현 caffe2
matheus695p/regularization-techniques-pytorch pytorch

Similar Papers 제목 키워드 기반

Cyclical Focal Loss

2022-02-16 · Leslie N. Smith

The cross-entropy softmax loss is the primary loss function used to train deep neural networks. On the other hand, the focal loss function has been demonstrated to provide improved performance when there is an imbalance …

General Cyclical Training of Neural Networks

2022-02-17 · Leslie N. Smith

This paper describes the principle of "General Cyclical Training" in machine learning, where training starts and ends with "easy training" and the "hard training" happens during the middle epochs. We propose several mani…

Data AugmentationKnowledge Distillation

Applying Cyclical Learning Rate to Neural Machine Translation

2020-04-06 · Choon Meng Lee, Jianfeng Liu, Wei Peng

In training deep learning networks, the optimizer and related learning rate are often used without much thought or with minimal tuning, even though it is crucial in ensuring a fast convergence to a good quality minimum o…

Machine TranslationTranslation

Privacy of the last iterate in cyclically-sampled DP-SGD on nonconvex composite losses

2024-07-07 · Weiwei Kong, Mónica Ribero

Differentially-private stochastic gradient descent (DP-SGD) is a family of iterative machine learning training algorithms that privatize gradients to generate a sequence of differentially-private (DP) model parameters. I…

Basel III capital surcharges for G-SIBs fail to control systemic risk and can cause pro-cyclical side effects

2016-02-10

In addition to constraining bilateral exposures of financial institutions, there are essentially two options for future financial regulation of systemic risk (SR): First, financial regulation could attempt to reduce the …