paper-with-me

홈 › Papers

The Break-Even Point on Optimization Trajectories of Deep Neural Networks

2020-02-21 · ICLR 2020 1 · Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, Krzysztof Geras

The early phase of training of deep neural networks is critical for their final performance. In this work, we study how the hyperparameters of stochastic gradient descent (SGD) used in the early phase of training affect the rest of the optimization trajectory. We argue for the existence of the "break-even" point on this trajectory, beyond which the curvature of the loss surface and noise in the gradient are implicitly regularized by SGD. In particular, we demonstrate on multiple classification tasks that using a large learning rate in the initial phase of training reduces the variance of the gradient, and improves the conditioning of the covariance of gradients. These effects are beneficial from the optimization perspective and become visible after the break-even point. Complementing prior work, we also show that using a low learning rate results in bad conditioning of the loss surface even for a neural network with batch normalization layers. In short, our work shows that key properties of the loss surface are strongly influenced by SGD in the early phase of training. We argue that studying the impact of the identified effects on generalization is a promising future direction.

📄 PDF Abstract BibTeX arXiv:2002.09572

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Batch Normalization 설명 없음
SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models

2026-03-11 · Xiangwen Wang, Ananth Balashankar, Varun Chandrasekaran arxiv

Large language models remain vulnerable to jailbreak attacks, yet we still lack a systematic understanding of how jailbreak success scales with attacker effort across methods, model families, and harm types. We initiate …

Doob's Lagrangian: A Sample-Efficient Variational Approach to Transition Path Sampling

2024-10-10 · Yuanqi Du, Michael Plainer, Rob Brekelmans, Chenru Duan 외

Rare event sampling in dynamical systems is a fundamental problem arising in the natural sciences, which poses significant computational challenges due to an exponentially large space of trajectories. For settings where …

Protein Folding

Spatiotemporal Transformers for Predicting Avian Disease Risk from Migration Trajectories

2025-10-17 · Dingya Feng, Dingyuan Xue arxiv

Accurate forecasting of avian disease outbreaks is critical for wildlife conservation and public health. This study presents a Transformer-based framework for predicting the disease risk at the terminal locations of migr…

Near-Future Policy Optimization

2026-04-22 · Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerates RLVR convergence and raises the performanc…

Reinforcement Learning

Learning the Intrinsic Dimensionality of Fermi-Pasta-Ulam-Tsingou Trajectories: A Nonlinear Approach using a Deep Autoencoder Model

2026-01-27 · Gionni Marchetti arxiv

We address the intrinsic dimensionality (ID) of high-dimensional trajectories, comprising $n_s = 4\,000\,000$ data points, of the Fermi-Pasta-Ulam-Tsingou (FPUT) $β$ model with $N = 32$ oscillators. To this end, a deep a…