paper-with-me

홈 › Papers

HammingMesh: A Network Topology for Large-Scale Deep Learning

2022-09-03 · Torsten Hoefler, Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo, Shigang Li, Marco Heddes, Jon Belk, Deepak Goel, Miguel Castro, Steve Scott

Numerous microarchitectural optimizations unlocked tremendous processing power for deep neural networks that in turn fueled the AI revolution. With the exhaustion of such optimizations, the growth of modern AI is now gated by the performance of training systems, especially their data movement. Instead of focusing on single accelerators, we investigate data-movement characteristics of large-scale training at full system scale. Based on our workload analysis, we design HammingMesh, a novel network topology that provides high bandwidth at low cost with high job scheduling flexibility. Specifically, HammingMesh can support full bandwidth and isolation to deep learning training jobs with two dimensions of parallelism. Furthermore, it also supports high global bandwidth for generic traffic. Thus, HammingMesh will power future large-scale deep learning systems with extreme bandwidth requirements.

📄 PDF Abstract BibTeX arXiv:2209.01346

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningScheduling

Similar Papers 제목 키워드 기반

Cover Learning for Large-Scale Topology Representation

2025-03-12 · Luis Scoccola, Uzu Lim, Heather A. Harrington

Classical unsupervised learning methods like clustering and linear dimensionality reduction parametrize large-scale geometry when it is discrete or linear, while more modern methods from manifold learning find low dimens…

Dimensionality ReductionTopological Data Analysis

FibeRed: Fiberwise Dimensionality Reduction of Topologically Complex Data with Vector Bundles

2022-06-13 · Luis Scoccola, Jose A. Perea

Datasets with non-trivial large scale topology can be hard to embed in low-dimensional Euclidean space with existing dimensionality reduction algorithms. We propose to model topologically complex datasets using vector bu…

Dimensionality Reduction

THD-BAR: Topology Hierarchical Derived Brain Autoregressive Modeling for EEG Generic Representations

2025-11-05 · Wenchao Yang, Weidong Yan, Wenkang Liu, Yulan Ma 외 arxiv

Large-scale pre-trained models hold significant potential for learning universal EEG representations. However, most existing methods, particularly autoregressive (AR) frameworks, primarily rely on straightforward tempora…

TA-MoE: Topology-Aware Large Scale Mixture-of-Expert Training

2023-02-20 · Chang Chen, Min Li, Zhihua Wu, dianhai yu 외

Sparsely gated Mixture-of-Expert (MoE) has demonstrated its effectiveness in scaling up deep neural networks to an extreme scale. Despite that numerous efforts have been made to improve the performance of MoE from the mo…

Data-driven Modeling of Linearizable Power Flow for Large-scale Grid Topology Optimization

2024-09-21 · Young-ho Cho, Hao Zhu

Effective power flow (PF) modeling critically affects the solution accuracy and computational complexity of large-scale grid optimization problems. Especially for grid optimization involving flexible topology to enhance …

Computational Efficiency