paper-with-me

홈 › Papers

ByzShield: An Efficient and Robust System for Distributed Training

2020-10-10 · Konstantinos Konstantinidis, Aditya Ramamoorthy

Training of large scale models on distributed clusters is a critical component of the machine learning pipeline. However, this training can easily be made to fail if some workers behave in an adversarial (Byzantine) fashion whereby they return arbitrary results to the parameter server (PS). A plethora of existing papers consider a variety of attack models and propose robust aggregation and/or computational redundancy to alleviate the effects of these attacks. In this work we consider an omniscient attack model where the adversary has full knowledge about the gradient computation assignments of the workers and can choose to attack (up to) any q out of K worker nodes to induce maximal damage. Our redundancy-based method ByzShield leverages the properties of bipartite expander graphs for the assignment of tasks to workers; this helps to effectively mitigate the effect of the Byzantine behavior. Specifically, we demonstrate an upper bound on the worst case fraction of corrupted gradients based on the eigenvalues of our constructions which are based on mutually orthogonal Latin squares and Ramanujan graphs. Our numerical experiments indicate over a 36% reduction on average in the fraction of corrupted gradients compared to the state of the art. Likewise, our experiments on training followed by image classification on the CIFAR-10 dataset show that ByzShield has on average a 20% advantage in accuracy under the most sophisticated attacks. ByzShield also tolerates a much larger fraction of adversarial nodes compared to prior work.

📄 PDF Abstract BibTeX arXiv:2010.04902

Code (1)

kkonstantinidis/ByzShield 공식 구현 pytorch

Tasks

image-classificationImage Classification

Similar Papers 제목 키워드 기반

Graph Neural Network-Based Distributed Optimal Control for Linear Networked Systems: An Online Distributed Training Approach

2025-04-08 · Zihao Song, Panos J. Antsaklis, Hai Lin

In this paper, we consider the distributed optimal control problem for linear networked systems. In particular, we are interested in learning distributed optimal controllers using graph recurrent neural networks (GRNNs).…

Distributed OptimizationGraph Neural NetworkSelf-Supervised Learning

Distributed Graph Neural Network Training: A Survey

2022-11-01 · Yingxia Shao, Hongzheng Li, Xizhi Gu, Hongbo Yin 외

Graph neural networks (GNNs) are a type of deep learning models that are trained on graphs and have been successfully applied in various domains. Despite the effectiveness of GNNs, it is still challenging for GNNs to eff…

CPUDistributed ComputingGPUGraph Neural Network+1

ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale

2023-03-24 · William Won, Taekyung Heo, Saeed Rashidi, Srinivas Sridharan 외

As deep learning models and input data are scaling at an unprecedented rate, it is inevitable to move towards distributed training platforms to fit the model and increase training throughput. State-of-the-art approaches …

HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models

2024-05-25 · Si Xu, Zixiao Huang, Yan Zeng, Shengen Yan 외

Training large-scale models relies on a vast number of computing resources. For example, training the GPT-4 model (1.8 trillion parameters) requires 25000 A100 GPUs . It is a challenge to build a large-scale cluster with…

GPU

Towards Scalable Distributed Training of Deep Learning on Public Cloud Clusters

2020-10-20 · Shaohuai Shi, Xianhao Zhou, Shutao Song, Xingyao Wang 외

Distributed training techniques have been widely deployed in large-scale deep neural networks (DNNs) training on dense-GPU clusters. However, on public cloud clusters, due to the moderate inter-connection bandwidth betwe…

GPU