paper-with-me

홈 › Papers

Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity

2025-08-07 · Hikaru Umeda, Hideaki Iiduka arxiv

The unprecedented growth of deep learning models has enabled remarkable advances but introduced substantial computational bottlenecks. A key factor contributing to training efficiency is batch-size and learning-rate scheduling in stochastic gradient methods. However, naive scheduling of these hyperparameters can degrade optimization efficiency and compromise generalization. Motivated by recent theoretical insights, we investigated how the batch size and learning rate should be increased during training to balance efficiency and convergence. We analyzed this problem on the basis of stochastic first-order oracle (SFO) complexity, defined as the expected number of gradient evaluations needed to reach an $ε$-approximate stationary point of the empirical loss. We theoretically derived optimal growth schedules for the batch size and learning rate that reduce SFO complexity and validated them through extensive experiments. Our results offer both theoretical insights and practical guidelines for scalable and efficient large-batch training in deep learning.

📄 PDF Abstract BibTeX arXiv:2508.05297

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards joint scaling laws with optimal batch size schedules

2026-07-30 · Jiaxiang Li, Zhiqi Bu, Shiyun Xu arxiv

Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics. In this paper, we study the deep learning dynami…

Unlocking optimal batch size schedules using continuous-time control and perturbation theory

2023-12-04 · Stefan Perko

Stochastic Gradient Descent (SGD) and its variants are almost universally used to train neural networks and to fit a variety of other parametric models. An important hyperparameter in this context is the batch size, whic…

Select without Fear: Almost All Mini-Batch Schedules Generalize Optimally

2023-05-03 · Konstantinos E. Nikolakakis, Amin Karbasi, Dionysis Kalogerias

We establish matching upper and lower generalization error bounds for mini-batch Gradient Descent (GD) training with either deterministic or stochastic, data-independent, but otherwise arbitrary batch selection rules. We…

All

Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling Laws

2026-02-15 · Jinbo Wang, Binghui Li, Zhanpeng Zhou, Mingze Wang 외 arxiv

Batch size scheduling (BSS) plays a critical role in large-scale deep learning training, influencing both optimization dynamics and computational efficiency. Yet, its theoretical foundations remain poorly understood. In …

Computational Efficiency

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism

2024-12-30 · Tim Tsz-Kit Lau, Weijian Li, Chenwei Xu, Han Liu 외

An appropriate choice of batch sizes in large-scale model training is crucial, yet it involves an intrinsic yet inevitable dilemma: large-batch training improves training efficiency in terms of memory utilization, while …