paper-with-me

홈 › Papers

Exploring Landscapes for Better Minima along Valleys

2025-10-31 · Tong Zhao, Jiacheng Li, Yuanchang Zhou, Guangming Tan, Weile Jia arxiv

Finding lower and better-generalizing minima is crucial for deep learning. However, most existing optimizers stop searching the parameter space once they reach a local minimum. Given the complex geometric properties of the loss landscape, it is difficult to guarantee that such a point is the lowest or provides the best generalization. To address this, we propose an adaptor "E" for gradient-based optimizers. The adapted optimizer tends to continue exploring along landscape valleys (areas with low and nearly identical losses) in order to search for potentially better local minima even after reaching a local minimum. This approach increases the likelihood of finding a lower and flatter local minimum, which is often associated with better generalization. We also provide a proof of convergence for the adapted optimizers in both convex and non-convex scenarios for completeness. Finally, we demonstrate their effectiveness in an important but notoriously difficult training scenario, large-batch training, where Lamb is the benchmark optimizer. Our testing results show that the adapted Lamb, ALTO, increases the test accuracy (generalization) of the current state-of-the-art optimizer by an average of 2.5% across a variety of large-batch training tasks. This work potentially opens a new research direction in the design of optimization algorithms.

📄 PDF Abstract BibTeX arXiv:2510.27153

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Asymmetric Valleys: Beyond Sharp and Flat Local Minima

2019-02-02 · NeurIPS 2019 12 · Haowei He, Gao Huang, Yang Yuan

Despite the non-convex nature of their loss functions, deep neural networks are known to generalize well when optimized with stochastic gradient descent (SGD). Recent work conjectures that SGD with proper configuration i…

Structure of the space of folding protein sequences defined by large language models

2023-11-10 · A. Zambon, R. Zecchina, G. Tiana

Proteins populate a manifold in the high-dimensional sequence space whose geometrical structure guides their natural evolution. Leveraging recently-developed structure prediction tools based on transformer models, we fir…

Symmetries, flat minima, and the conserved quantities of gradient flow

2022-10-31 · Bo Zhao, Iordan Ganev, Robin Walters, Rose Yu 외

Empirical studies of the loss landscape of deep networks have revealed that many local minima are connected through low-loss valleys. Yet, little is known about the theoretical origin of such valleys. We present a genera…

Spurious Valleys in Two-layer Neural Network Optimization Landscapes

2018-02-18 · Luca Venturi, Afonso S. Bandeira, Joan Bruna

Neural networks provide a rich class of high-dimensional, non-convex optimization problems. Despite their non-convexity, gradient-descent methods often successfully optimize these models. This has motivated a recent spur…

Vocal Bursts Valence Prediction

Getting higher on rugged landscapes: Inversion mutations open access to fitter adaptive peaks in NK fitness landscapes

2022-10-13 · Leonardo Trujillo, Paul Banse, Guillaume Beslon

Molecular evolution is often conceptualised as adaptive walks on rugged fitness landscapes, driven by mutations and constrained by incremental fitness selection. It is well known that epistasis shapes the ruggedness of t…