paper-with-me

홈 › Papers

Two Sides of One Coin: the Limits of Untuned SGD and the Power of Adaptive Methods

2023-05-21 · NeurIPS 2023 11

The classical analysis of Stochastic Gradient Descent (SGD) with polynomially decaying stepsize $\eta_t = \eta/\sqrt{t}$ relies on well-tuned $\eta$ depending on problem parameters such as Lipschitz smoothness constant, which is often unknown in practice. In this work, we prove that SGD with arbitrary $\eta > 0$, referred to as untuned SGD, still attains an order-optimal convergence rate $\widetilde{O}(T^{-1/4})$ in terms of gradient norm for minimizing smooth objectives. Unfortunately, it comes at the expense of a catastrophic exponential dependence on the smoothness constant, which we show is unavoidable for this scheme even in the noiseless setting. We then examine three families of adaptive methods $\unicode{x2013}$ Normalized SGD (NSGD), AMSGrad, and AdaGrad $\unicode{x2013}$ unveiling their power in preventing such exponential dependency in the absence of information about the smoothness parameter and boundedness of stochastic gradients. Our results provide theoretical justification for the advantage of adaptive methods over untuned SGD in alleviating the issue with large gradients.

📄 PDF Abstract BibTeX arXiv:2305.12475

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

AMSGrad AMSGrad is a stochastic optimization method that seeks to fix a convergence issue with Adam based optimizers. AMSGrad uses the…
AdaGrad AdaGrad is a stochastic optimization method that adapts the learning rate to the parameters. It performs smaller updates for parameters associated with frequently occurring…
SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Quantum Kernel Advantage over Classical Collapse in Medical Foundation Model Embeddings

2026-04-27 · Sebastian Cajas Ordóñez, Felipe Ocampo Osorio, Dax Enshan Koh, Rafi Al Attrach 외 arxiv

We provide evidence of quantum kernel advantage under noiseless simulation in binary insurance classification on MIMIC-CXR chest radiographs using quantum support vector machines (QSVM) with frozen embeddings from three …

On the adequacy of untuned warmup for adaptive optimization

2019-10-09 · Jerry Ma, Denis Yarats

Adaptive optimization algorithms such as Adam are widely used in deep learning. The stability of such algorithms is often improved with a warmup schedule for the learning rate. Motivated by the difficulty of choosing and…

Image ClassificationLanguage ModellingMachine Translation

The Power of Adaptivity in Identifying Statistical Alternatives

2016-12-01 · NeurIPS 2016 12 · Kevin G. Jamieson, Daniel Haas, Benjamin Recht

This paper studies the trade-off between two different kinds of pure exploration: breadth versus depth. We focus on the most biased coin problem, asking how many total coin flips are required to identify a ``heavy'' coin…

Anomaly Detection

LEAF: Unveiling Two Sides of the Same Coin in Semi-supervised Facial Expression Recognition

2024-04-23 · Fan Zhang, Zhi-Qi Cheng, Jian Zhao, Xiaojiang Peng 외

Semi-supervised learning has emerged as a promising approach to tackle the challenge of label scarcity in facial expression recognition (FER) task. However, current state-of-the-art methods primarily focus on one side of…

Facial Expression RecognitionFacial Expression Recognition (FER)

Automated Coin Recognition System using ANN

2013-12-23 · Shatrughan Modi, Dr. Seema Bawa

Coins are integral part of our day to day life. We use coins everywhere like grocery store, banks, buses, trains etc. So it becomes a basic need that coins can be sorted and counted automatically. For this it is necessar…