paper-with-me

홈 › Papers

Delayed Gradient Averaging: Tolerate the Communication Latency for Federated Learning

2021-12-01 · NeurIPS 2021 12 · Ligeng Zhu, Hongzhou Lin, Yao Lu, Yujun Lin, Song Han

Federated Learning is an emerging direction in distributed machine learning that en-ables jointly training a model without sharing the data. Since the data is distributed across many edge devices through wireless / long-distance connections, federated learning suffers from inevitable high communication latency. However, the latency issues are undermined in the current literature [15] and existing approaches suchas FedAvg [27] become less efficient when the latency increases. To over comethe problem, we propose \textbf{D}elayed \textbf{G}radient \textbf{A}veraging (DGA), which delays the averaging step to improve efficiency and allows local computation in parallel tocommunication. We theoretically prove that DGA attains a similar convergence rate as FedAvg, and empirically show that our algorithm can tolerate high network latency without compromising accuracy. Specifically, we benchmark the training speed on various vision (CIFAR, ImageNet) and language tasks (Shakespeare),with both IID and non-IID partitions, and show DGA can bring 2.55$\times$ to 4.07$\times$ speedup. Moreover, we built a 16-node Raspberry Pi cluster and show that DGA can consistently speed up real-world federated learning applications.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Federated Learning

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

DaSGD: Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging

2020-05-31 · Qinggang Zhou, Yawen Zhang, Pengcheng Li, Xiaoyong Liu 외

The state-of-the-art deep learning algorithms rely on distributed training systems to tackle the increasing sizes of models and training data sets. Minibatch stochastic gradient descent (SGD) algorithm requires workers t…

Delayed Random Partial Gradient Averaging for Federated Learning

2024-12-28 · Xinyi Hu

Federated learning (FL) is a distributed machine learning paradigm that enables multiple clients to train a shared model collaboratively while preserving privacy. However, the scaling of real-world FL systems is often li…

Federated Learning

Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging

2021-09-29 · Pengcheng Li, Yixin Guo, Yawen Zhang, Qinggang Zhou

State-of-the-art deep learning algorithms rely on distributed training to tackle the increasing model size and training data. Mini-batch Stochastic Gradient Descent (SGD) requires workers to halt forward/backward propaga…

Smart Split-Federated Learning over Noisy Channels for Embryo Image Segmentation

2026-01-26 · Zahra Hafezi Kafshgari, Ivan V. Bajic, Parvaneh Saeedi arxiv

Split-Federated (SplitFed) learning is an extension of federated learning that places minimal requirements on the clients computing infrastructure, since only a small portion of the overall model is deployed on the clien…

Federated LearningImage Segmentation

DBLP: Phase-Aware Bounded-Loss Transport for Burst-Resilient Distributed ML Training

2026-05-03 · Zechen Ma, Zixi Qu, Jinyan Yi, David Lin 외 arxiv

Distributed machine learning (ML) training has become a necessity with the prevalence of billion to trillion-parameter-scale models. While prior work has improved training efficiency from the ML perspective at the applic…