HammingMesh: A Network Topology for Large-Scale Deep Learning
Numerous microarchitectural optimizations unlocked tremendous processing power for deep neural networks that in turn fueled the AI revolution. With the exhaustion of such optimizations, the growth of modern AI is now gated by the performance of training systems, especially their data movement. Instead of focusing on single accelerators, we investigate data-movement characteristics of large-scale training at full system scale. Based on our workload analysis, we design HammingMesh, a novel network topology that provides high bandwidth at low cost with high job scheduling flexibility. Specifically, HammingMesh can support full bandwidth and isolation to deep learning training jobs with two dimensions of parallelism. Furthermore, it also supports high global bandwidth for generic traffic. Thus, HammingMesh will power future large-scale deep learning systems with extreme bandwidth requirements.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningSchedulingSimilar Papers 제목 키워드 기반
Cover Learning for Large-Scale Topology Representation
Classical unsupervised learning methods like clustering and linear dimensionality reduction parametrize large-scale geometry when it is discrete or linear, while more modern methods from manifold learning find low dimens…
Dimensionality ReductionTopological Data AnalysisFibeRed: Fiberwise Dimensionality Reduction of Topologically Complex Data with Vector Bundles
Datasets with non-trivial large scale topology can be hard to embed in low-dimensional Euclidean space with existing dimensionality reduction algorithms. We propose to model topologically complex datasets using vector bu…
Dimensionality ReductionTHD-BAR: Topology Hierarchical Derived Brain Autoregressive Modeling for EEG Generic Representations
Large-scale pre-trained models hold significant potential for learning universal EEG representations. However, most existing methods, particularly autoregressive (AR) frameworks, primarily rely on straightforward tempora…
TA-MoE: Topology-Aware Large Scale Mixture-of-Expert Training
Sparsely gated Mixture-of-Expert (MoE) has demonstrated its effectiveness in scaling up deep neural networks to an extreme scale. Despite that numerous efforts have been made to improve the performance of MoE from the mo…
Data-driven Modeling of Linearizable Power Flow for Large-scale Grid Topology Optimization
Effective power flow (PF) modeling critically affects the solution accuracy and computational complexity of large-scale grid optimization problems. Especially for grid optimization involving flexible topology to enhance …
Computational Efficiency