paper-with-me

홈 › Papers

Better scalability under potentially heavy-tailed gradients

2020-06-01 · Matthew J. Holland

We study a scalable alternative to robust gradient descent (RGD) techniques that can be used when the gradients can be heavy-tailed, though this will be unknown to the learner. The core technique is simple: instead of trying to robustly aggregate gradients at each step, which is costly and leads to sub-optimal dimension dependence in risk bounds, we choose a candidate which does not diverge too far from the majority of cheap stochastic sub-processes run for a single pass over partitioned data. In addition to formal guarantees, we also provide empirical analysis of robustness to perturbations to experimental conditions, under both sub-Gaussian and heavy-tailed data. The result is a procedure that is simple to implement, trivial to parallelize, which keeps the formal strength of RGD methods but scales much better to large learning problems.

📄 PDF Abstract BibTeX arXiv:2006.00784

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Better scalability under potentially heavy-tailed feedback

2020-12-14 · Matthew J. Holland

We study scalable alternatives to robust gradient descent (RGD) techniques that can be used when the losses and/or gradients can be heavy-tailed, though this will be unknown to the learner. The core technique is simple: …

Learning with CVaR-based feedback under potentially heavy tails

2020-06-03 · Matthew J. Holland, El Mehdi Haress

We study learning algorithms that seek to minimize the conditional value-at-risk (CVaR), when all the learner knows is that the losses incurred may be heavy-tailed. We begin by studying a general-purpose estimator of CVa…

Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise

2025-02-02 · Thanh Dang, Melih Barsbey, A K M Rokonuzzaman Sonet, Mert Gurbuzbalaban 외

Understanding the generalization properties of optimization algorithms under heavy-tailed noise has gained growing attention. However, the existing theoretical results mainly focus on stochastic gradient descent (SGD) an…

Generalization Bounds

Multi-agent Multi-armed Bandit with Fully Heavy-tailed Dynamics

2025-01-31 · Xingyu Wang, Mengfan Xu

We study decentralized multi-agent multi-armed bandits in fully heavy-tailed settings, where clients communicate over sparse random graphs with heavy-tailed degree distributions and observe heavy-tailed (homogeneous or h…

Multi-Armed Bandits

Convergence Rates of Stochastic Gradient Descent under Infinite Noise Variance

2021-02-20 · NeurIPS 2021 12 · Hongjian Wang, Mert Gürbüzbalaban, Lingjiong Zhu, Umut Şimşekli 외

Recent studies have provided both empirical and theoretical evidence illustrating that heavy tails can emerge in stochastic gradient descent (SGD) in various scenarios. Such heavy tails potentially result in iterates wit…