Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
It has been experimentally observed that distributed implementations of mini-batch stochastic gradient descent (SGD) algorithms exhibit speedup saturation and decaying generalization ability beyond a particular batch-size. In this work, we present an analysis hinting that high similarity between concurrently processed gradients may be a cause of this performance degradation. We introduce the notion of gradient diversity that measures the dissimilarity between concurrent gradient updates, and show its key role in the performance of mini-batch SGD. We prove that on problems with high gradient diversity, mini-batch SGD is amenable to better speedups, while maintaining the generalization performance of serial (one sample) SGD. We further establish lower bounds on convergence where mini-batch SGD slows down beyond a particular batch-size, solely due to the lack of gradient diversity. We provide experimental evidence indicating the key role of gradient diversity in distributed learning, and discuss how heuristics like dropout, Langevin dynamics, and quantization can improve it.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityQuantizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
How should fishing mortality be distributed under balanced harvesting?
Zhou and Smith (2017) investigate different multi-species harvesting scenarios using a simple Holling-Tanner model. Among these scenarios are two methods for implementing balanced harvesting, where fishing is distributed…
High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise
High-probability analysis of stochastic first-order optimization methods under mild assumptions on the noise has been gaining a lot of attention in recent years. Typically, gradient clipping is one of the key algorithmic…
Distributed OptimizationScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed Training
Large-scale distributed training of Deep Neural Networks (DNNs) on state-of-the-art platforms is expected to be severely communication constrained. To overcome this limitation, numerous gradient compression techniques ha…
MobileCLIP2: Improving Multi-Modal Reinforced Training
Foundation image-text models such as CLIP with zero-shot capabilities enable a wide array of applications. MobileCLIP is a recent family of image-text models at 3-15ms latency and 50-150M parameters with state-of-the-art…
Knowledge DistillationQoS-aware Big Service Composition using Distributed Co-Evolutionary Algorithm
Big services are collections of interrelated web services across virtual and physical domains, processing Big Data. Existing service selection and composition algorithms fail to achieve the global optimum solution in a r…
DiversityMultiobjective OptimizationService Composition