paper-with-me

홈 › Papers

Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates

2019-11-20 · Cong Xie, Oluwasanmi Koyejo, Indranil Gupta, Haibin Lin

When scaling distributed training, the communication overhead is often the bottleneck. In this paper, we propose a novel SGD variant with reduced communication and adaptive learning rates. We prove the convergence of the proposed algorithm for smooth but non-convex problems. Empirical results show that the proposed algorithm significantly reduces the communication overhead, which, in turn, reduces the training time by up to 30% for the 1B word dataset.

📄 PDF Abstract BibTeX arXiv:1911.09030

Code (1)

xcgoner/AISTATS2020-AdaAlter-GluonNLP 공식 구현 mxnet

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Local Stochastic Gradient Descent Ascent: Convergence Analysis and Communication Efficiency

2021-02-25 · Yuyang Deng, Mehrdad Mahdavi

Local SGD is a promising approach to overcome the communication overhead in distributed learning by reducing the synchronization frequency among worker nodes. Despite the recent theoretical advances of local SGD in empir…

Communication-Efficient Distributed SGD with Compressed Sensing

2021-12-15 · Yujie Tang, Vikram Ramanathan, Junshan Zhang, Na Li

We consider large scale distributed optimization over a set of edge devices connected to a central server, where the limited communication bandwidth between the server and edge devices imposes a significant bottleneck fo…

compressed sensingDistributed OptimizationFederated Learning

A Communication Efficient Collaborative Learning Framework for Distributed Features

2019-12-24 · Yang Liu, Yan Kang, Xinwei Zhang, Liping Li 외

We introduce a collaborative learning framework allowing multiple parties having different sets of attributes about the same user to jointly build models without exposing their raw data or model parameters. In particular…

STL-SGD: Speeding Up Local SGD with Stagewise Communication Period

2020-06-11 · Shuheng Shen, Yifei Cheng, Jingchang Liu, Linli Xu

Distributed parallel stochastic gradient descent algorithms are workhorses for large scale machine learning tasks. Among them, local stochastic gradient descent (Local SGD) has attracted significant attention due to its …

Flattened one-bit stochastic gradient descent: compressed distributed optimization with controlled variance

2024-05-17 · Alexander Stollenwerk, Laurent Jacques

We propose a novel algorithm for distributed stochastic gradient descent (SGD) with compressed gradient communication in the parameter-server framework. Our gradient compression technique, named flattened one-bit stochas…

Distributed OptimizationQuantization