paper-with-me

홈 › Papers

Error Feedback Reloaded: From Quadratic to Arithmetic Mean of Smoothness Constants

2024-02-16 · Peter Richtárik, Elnur Gasanov, Konstantin Burlachenko

Error Feedback (EF) is a highly popular and immensely effective mechanism for fixing convergence issues which arise in distributed training methods (such as distributed GD or SGD) when these are enhanced with greedy communication compression techniques such as TopK. While EF was proposed almost a decade ago (Seide et al., 2014), and despite concentrated effort by the community to advance the theoretical understanding of this mechanism, there is still a lot to explore. In this work we study a modern form of error feedback called EF21 (Richtarik et al., 2021) which offers the currently best-known theoretical guarantees, under the weakest assumptions, and also works well in practice. In particular, while the theoretical communication complexity of EF21 depends on the quadratic mean of certain smoothness parameters, we improve this dependence to their arithmetic mean, which is always smaller, and can be substantially smaller, especially in heterogeneous data regimes. We take the reader on a journey of our discovery process. Starting with the idea of applying EF21 to an equivalent reformulation of the underlying problem which (unfortunately) requires (often impractical) machine cloning, we continue to the discovery of a new weighted version of EF21 which can (fortunately) be executed without any cloning, and finally circle back to an improved analysis of the original EF21 method. While this development applies to the simplest form of EF21, our approach naturally extends to more elaborate variants involving stochastic gradients and partial participation. Further, our technique improves the best-known theory of EF21 in the rare features regime (Richtarik et al., 2023). Finally, we validate our theoretical findings with suitable experiments.

📄 PDF Abstract BibTeX arXiv:2402.10774

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Optimal Receive Beamforming for Over-the-Air Computation

2021-05-11 · Wenzhi Fang, Yinan Zou, Hongbin Zhu, Yuanming Shi 외

In this paper, we consider fast wireless data aggregation via over-the-air computation (AirComp) in Internet of Things (IoT) networks, where an access point (AP) with multiple antennas aim to recover the arithmetic mean …

Denoising

Simple Modulo can Significantly Outperform Deep Learning-based Deepcode

2020-08-04 · Assaf Ben-Yishai, Ofer Shayevitz

Deepcode (H.Kim et al.2018) is a recently suggested Deep Learning-based scheme for communication over the AWGN channel with noisy feedback, claimed to be superior to all previous schemes in the literature. Deepcode's use…

Deep Learning

On the Constructing Bifurcation Diagram of the Quadratic Map With Floating-Point Arithmetic

2017-11-28

This paper presents an analysis on the effects of floating-point arithmetic on the constructing bifurcation diagram of the quadratic map. More precisely, we are interested in showing the dependence of initial conditions …

Decentralized Non-convex Stochastic Optimization with Heterogeneous Variance

2026-02-12 · Hongxu Chen, Ke Wei, Luo Luo arxiv

Decentralized optimization is critical for solving large-scale machine learning problems over distributed networks, where multiple nodes collaborate through local communication. In practice, the variances of stochastic g…

Stochastic Optimization

What do you Mean? The Role of the Mean Function in Bayesian Optimisation

2020-04-17 · George De Ath, Jonathan E. Fieldsend, Richard M. Everson

Bayesian optimisation is a popular approach for optimising expensive black-box functions. The next location to be evaluated is selected via maximising an acquisition function that balances exploitation and exploration. G…

Bayesian OptimisationGaussian Processes