paper-with-me

홈 › Papers

Demystifying Why Local Aggregation Helps: Convergence Analysis of Hierarchical SGD

2020-10-24 · Jiayi Wang, Shiqiang Wang, Rong-Rong Chen, Mingyue Ji

Hierarchical SGD (H-SGD) has emerged as a new distributed SGD algorithm for multi-level communication networks. In H-SGD, before each global aggregation, workers send their updated local models to local servers for aggregations. Despite recent research efforts, the effect of local aggregation on global convergence still lacks theoretical understanding. In this work, we first introduce a new notion of "upward" and "downward" divergences. We then use it to conduct a novel analysis to obtain a worst-case convergence upper bound for two-level H-SGD with non-IID data, non-convex objective function, and stochastic gradient. By extending this result to the case with random grouping, we observe that this convergence upper bound of H-SGD is between the upper bounds of two single-level local SGD settings, with the number of local iterations equal to the local and global update periods in H-SGD, respectively. We refer to this as the "sandwich behavior". Furthermore, we extend our analytical approach based on "upward" and "downward" divergences to study the convergence for the general case of H-SGD with more than two levels, where the "sandwich behavior" still holds. Our theoretical results provide key insights of why local aggregation can be beneficial in improving the convergence of H-SGD.

📄 PDF Abstract BibTeX arXiv:2010.12998

Code (1)

c3atuofu/hierarchical-sgd 공식 구현 pytorch

Tasks

Federated Learning

Similar Papers 제목 키워드 기반

Decentralized Federated Learning Over Imperfect Communication Channels

2024-05-21 · Weicai Li, Tiejun Lv, Wei Ni, Jingbo Zhao 외

This paper analyzes the impact of imperfect communication channels on decentralized federated learning (D-FL) and subsequently determines the optimal number of local aggregations per training round, adapting to the netwo…

Federated Learningimage-classificationImage Classification

Convergence Analysis of Aggregation-Broadcast in LoRA-enabled Distributed Fine-Tuning

2025-08-02 · Xin Chen, Shuaijun Chen, Omid Tavallaie, Nguyen Tran 외 arxiv

Federated Learning (FL) enables collaborative model training across decentralized data sources while preserving data privacy. However, the growing size of Machine Learning (ML) models poses communication and computation …

Federated Learning

Demystifying overcomplete nonlinear auto-encoders: fast SGD convergence towards sparse representation from random initialization

2018-01-01 · ICLR 2018 1 · Cheng Tang, Claire Monteleoni

Auto-encoders are commonly used for unsupervised representation learning and for pre-training deeper neural networks. When its activation function is linear and the encoding dimension (width of hidden layer) is smaller t…

Dictionary LearningRepresentation Learning

Over-the-Air Federated Learning and Optimization

2023-10-16 · Jingyang Zhu, Yuanming Shi, Yong Zhou, Chunxiao Jiang 외

Federated learning (FL), as an emerging distributed machine learning paradigm, allows a mass of edge devices to collaboratively train a global model while preserving privacy. In this tutorial, we focus on FL via over-the…

Federated Learning

Federated Adversarial Learning: A Framework with Convergence Analysis

2022-08-07 · Xiaoxiao Li, Zhao Song, Jiaming Yang

Federated learning (FL) is a trending training paradigm to utilize decentralized training data. FL allows clients to update model parameters locally for several epochs, then share them to a global model for aggregation. …

Federated Learning