paper-with-me

홈 › Papers

On the equivalence of different adaptive batch size selection strategies for stochastic gradient descent methods

2021-09-22 · Luis Espath, Sebastian Krumscheid, Raúl Tempone, Pedro Vilanova

In this study, we demonstrate that the norm test and inner product/orthogonality test presented in \cite{Bol18} are equivalent in terms of the convergence rates associated with Stochastic Gradient Descent (SGD) methods if $\epsilon^2=\theta^2+\nu^2$ with specific choices of $\theta$ and $\nu$. Here, $\epsilon$ controls the relative statistical error of the norm of the gradient while $\theta$ and $\nu$ control the relative statistical error of the gradient in the direction of the gradient and in the direction orthogonal to the gradient, respectively. Furthermore, we demonstrate that the inner product/orthogonality test can be as inexpensive as the norm test in the best case scenario if $\theta$ and $\nu$ are optimally selected, but the inner product/orthogonality test will never be more computationally affordable than the norm test if $\epsilon^2=\theta^2+\nu^2$. Finally, we present two stochastic optimization problems to illustrate our results.

📄 PDF Abstract BibTeX arXiv:2109.10933

Code (0)

등록된 구현이 없습니다.

Tasks

Stochastic Optimization

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Scalable Batch-Mode Deep Bayesian Active Learning via Equivalence Class Annealing

2021-12-27 · Renyu Zhang, Aly A. Khan, Robert L. Grossman, Yuxin Chen

Active learning has demonstrated data efficiency in many fields. Existing active learning algorithms, especially in the context of batch-mode deep Bayesian active models, rely heavily on the quality of uncertainty estima…

Active LearningDiversityMulti-class Classification

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling

2025-10-16 · Alexandru Meterez, Depen Morwani, Jingfeng Wu, Costin-Andrei Oncescu 외 arxiv

Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. While for SGD, doubling the batch size can be equivalent to halving the learning ra…

Big Batch SGD: Automated Inference using Adaptive Batch Sizes

2016-10-18 · Soham De, Abhay Yadav, David Jacobs, Tom Goldstein

Classical stochastic gradient methods for optimization rely on noisy gradient approximations that become progressively less accurate as iterates approach a solution. The large noise and small signal in the resulting grad…

Stochastic batch size for adaptive regularization in deep network optimization

2020-04-14 · Kensuke Nakamura, Stefano Soatto, Byung-Woo Hong

We propose a first-order stochastic optimization algorithm incorporating adaptive regularization applicable to machine learning problems in deep learning framework. The adaptive regularization is imposed by stochastic pr…

image-classificationImage ClassificationStochastic Optimization

Adaptive Batch Sizes for Active Learning A Probabilistic Numerics Approach

2023-06-09 · Masaki Adachi, Satoshi Hayakawa, Martin Jørgensen, Xingchen Wan 외

Active learning parallelization is widely used, but typically relies on fixing the batch size throughout experimentation. This fixed approach is inefficient because of a dynamic trade-off between cost and speed -- larger…

Active LearningBayesian OptimisationBayesian OptimizationDrug Discovery