paper-with-me

홈 › Papers

Benchmarking Optimizers for Large Language Model Pretraining

2025-09-01 · Andrei Semenov, Matteo Pagliardini, Martin Jaggi arxiv

The recent development of Large Language Models (LLMs) has been accompanied by an effervescence of novel ideas and methods to better optimize the loss of deep learning models. Claims from those methods are myriad: from faster convergence to removing reliance on certain hyperparameters. However, the diverse experimental protocols used to validate these claims make direct comparisons between methods challenging. This study presents a comprehensive evaluation of recent optimization techniques across standardized LLM pretraining scenarios, systematically varying model size, batch size, and training duration. Through careful tuning of each method, we provide guidance to practitioners on which optimizer is best suited for each scenario. For researchers, our work highlights promising directions for future optimization research. Finally, by releasing our code and making all experiments fully reproducible, we hope our efforts can help the development and rigorous benchmarking of future methods.

📄 PDF Abstract BibTeX arXiv:2509.01440

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers

2026-05-09 · Aditya Ranganath arxiv

Training large language models requires optimization algorithms that are not only statistically effective, but also computationally and memory efficient at extreme scale. Although Adam remains the dominant optimizer for …

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

2026-07-04 · Siyuan Li, Jiabao Pan, Yumou Liu, Zhuoli Ouyang 외 hf

Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning budget, and task diversity, yet the landscape of over one hundred methods remain…

Image Classification

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

2026-05-07 · Yuxing Liu, Jianyu Wang, Tong Zhang arxiv

Optimizers play an important role in both pretraining and finetuning stages when training large language models (LLMs). In this paper, we present an observation that full finetuning with the same optimizer as in pretrain…

When Do Flat Minima Optimizers Work?

2022-02-01 · Jean Kaddour, Linqing Liu, Ricardo Silva, Matt J. Kusner

Recently, flat-minima optimizers, which seek to find parameters in low-loss neighborhoods, have been shown to improve a neural network's generalization performance over stochastic and adaptive gradient-based optimizers. …

BenchmarkingGraph LearningGraph Representation LearningImage Classification+9

MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic Optimization

2026-06-16 · Da Chang, Ganzhao Yuan arxiv

Efficient optimization is essential for training large language models. Although intra-layer selective updates have been explored, a general mechanism that enables fine-grained control while ensuring convergence guarante…

Stochastic Optimization