paper-with-me

홈 › Papers

Pre-Training LLMs on a budget: A comparison of three optimizers

2025-07-11 · Joel Schlotthauer, Christian Kroos, Chris Hinze, Viktor Hangya, Luzian Hahn, Fabian Küch arxiv

Optimizers play a decisive role in reducing pre-training times for LLMs and achieving better-performing models. In this study, we compare three major variants: the de-facto standard AdamW, the simpler Lion, developed through an evolutionary search, and the second-order optimizer Sophia. For better generalization, we train with two different base architectures and use a single- and a multiple-epoch approach while keeping the number of tokens constant. Using the Maximal Update Parametrization and smaller proxy models, we tune relevant hyperparameters separately for each combination of base architecture and optimizer. We found that while the results from all three optimizers were in approximately the same range, Sophia exhibited the lowest training and validation loss, Lion was fastest in terms of training GPU hours but AdamW led to the best downstream evaluation results.

📄 PDF Abstract BibTeX arXiv:2507.08472

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fantastic Pretraining Optimizers and Where to Find Them

2025-09-02 · Kaiyue Wen, David Hall, Tengyu Ma, Percy Liang arxiv

AdamW has long been the dominant optimizer in language model pretraining, despite numerous claims that alternative optimizers offer 1.4 to 2x speedup. We posit that two methodological shortcomings have obscured fair comp…

Towards Robust Scaling Laws for Optimizers

2026-02-07 · Alexandra Volkova, Mher Safaryan, Christoph H. Lampert, Dan Alistarh arxiv

The quality of Large Language Model (LLM) pretraining depends on multiple factors, including the compute budget and the choice of optimization algorithm. Empirical scaling laws are widely used to predict loss as model si…

Meta-Learning for Black-box Optimization

2019-07-16 · Vishnu TV, Pankaj Malhotra, Jyoti Narwariya, Lovekesh Vig 외

Recently, neural networks trained as optimizers under the "learning to learn" or meta-learning framework have been shown to be effective for a broad range of optimization tasks including derivative-free black-box functio…

Meta-Learning

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers

2026-05-09 · Aditya Ranganath arxiv

Training large language models requires optimization algorithms that are not only statistically effective, but also computationally and memory efficient at extreme scale. Although Adam remains the dominant optimizer for …

MO-CAPO: Multi-Objective Cost-Aware Prompt Optimization

2026-05-15 · Jan Büssing, Moritz Schlager, Timo Heiß, Tom Zehle 외 arxiv

Large language models (LLMs) achieve strong performance across a wide range of tasks but are highly sensitive to prompt design, motivating the need for automatic prompt optimization. Existing methods predominantly focus …