paper-with-me

홈 › Papers

Error Feedback Fixes SignSGD and other Gradient Compression Schemes

2019-01-28 · Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian U. Stich, Martin Jaggi

Sign-based algorithms (e.g. signSGD) have been proposed as a biased gradient compression technique to alleviate the communication bottleneck in training large neural networks across multiple workers. We show simple convex counter-examples where signSGD does not converge to the optimum. Further, even when it does converge, signSGD may generalize poorly when compared with SGD. These issues arise because of the biased nature of the sign compression operator. We then show that using error-feedback, i.e. incorporating the error made by the compression operator into the next step, overcomes these issues. We prove that our algorithm EF-SGD with arbitrary compression operator achieves the same rate of convergence as SGD without any additional assumptions. Thus EF-SGD achieves gradient compression for free. Our experiments thoroughly substantiate the theory and show that error-feedback improves both convergence and generalization. Code can be found at \url{https://github.com/epfml/error-feedback-SGD}.

📄 PDF Abstract BibTeX arXiv:1901.09847

Code (2)

epfml/error-feedback-SGD 공식 구현 pytorch
MindSpore-scientific-2/code-11/tree/main/signSGD mindspore

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Magnitude Matters: Fixing SIGNSGD Through Magnitude-Aware Sparsification in the Presence of Data Heterogeneity

2023-02-19 · Richeng Jin, Xiaofan He, Caijun Zhong, Zhaoyang Zhang 외

Communication overhead has become one of the major bottlenecks in the distributed training of deep neural networks. To alleviate the concern, various gradient compression methods have been proposed, and sign-based algori…

Federated Learning

Compressing gradients in distributed SGD by exploiting their temporal correlation

2021-01-01 · Tharindu Adikari, Stark Draper

We propose SignXOR, a novel compression scheme that exploits temporal correlation of gradients for the purpose of gradient compression. Sign-based schemes such as Scaled-sign and SignSGD (Bernstein et al., 2018; Karimire…

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

2026-04-16 · Dingzhi Yu, Rui Pan, Yuxing Liu, Tong Zhang arxiv

Sign-based optimization algorithms, such as SignSGD, have garnered significant attention for their remarkable performance in distributed learning and training large foundation models. Despite their empirical superiority,…

Mathematical Reasoning

signSGD via Zeroth-Order Oracle

2019-05-01 · ICLR 2019 5 · Sijia Liu, Pin-Yu Chen, Xiangyi Chen, Mingyi Hong

In this paper, we design and analyze a new zeroth-order (ZO) stochastic optimization algorithm, ZO-signSGD, which enjoys dual advantages of gradient-free operations and signSGD. The latter requires only the sign informat…

image-classificationImage ClassificationStochastic Optimization

Robustness to Unbounded Smoothness of Generalized SignSGD

2022-08-23 · Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei zhang 외

Traditional analyses in non-convex optimization typically rely on the smoothness assumption, namely requiring the gradients to be Lipschitz. However, recent evidence shows that this smoothness condition does not capture …