paper-with-me

홈 › Papers

Scaling Deep Learning Training with MPMD Pipeline Parallelism

2024-12-18 · Anxhelo Xhebraj, Sean Lee, Hanfeng Chen, Vinod Grover

We present JaxPP, a system for efficiently scaling the training of large deep learning models with flexible pipeline parallelism. We introduce a seamless programming model that allows implementing user-defined pipeline schedules for gradient accumulation. JaxPP automatically distributes tasks, corresponding to pipeline stages, over a cluster of nodes and automatically infers the communication among them. We implement a MPMD runtime for asynchronous execution of SPMD tasks. The pipeline parallelism implementation of JaxPP improves hardware utilization by up to $1.11\times$ with respect to the best performing SPMD configuration.

📄 PDF Abstract BibTeX arXiv:2412.14374

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Learning

Similar Papers 제목 키워드 기반

The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs with Hybrid Parallelism

2020-07-25 · Yosuke Oyama, Naoya Maruyama, Nikoli Dryden, Erin McCarthy 외

We present scalable hybrid-parallel algorithms for training large-scale 3D convolutional neural networks. Deep learning-based emerging scientific workflows often require model training with large, high-dimensional sample…

2k

Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

2021-04-09 · Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick Legresley 외

Large language models have led to state-of-the-art accuracies across a range of tasks. However, training these models efficiently is challenging for two reasons: a) GPU memory capacity is limited, making it impossible to…

GPULanguage ModelingLanguage Modelling

AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism

2026-01-30 · Thalaiyasingam Ajanthan, Sameera Ramasinghe, Gil Avraham, Hadi Mohaghegh Dolatabadi 외 arxiv

Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limit…

GNNPipe: Scaling Deep GNN Training with Pipelined Model Parallelism

2023-08-19 · Jingji Chen, Zhuoming Chen, Xuehai Qian

Communication is a key bottleneck for distributed graph neural network (GNN) training. This paper proposes GNNPipe, a new approach that scales the distributed full-graph deep GNN training. Being the first to use layer-le…

GPUGraph Neural Network

Merak: An Efficient Distributed DNN Training Framework with Automated 3D Parallelism for Giant Foundation Models

2022-06-10 · Zhiquan Lai, Shengwei Li, Xudong Tang, Keshi Ge 외

Foundation models are becoming the dominant deep learning technologies. Pretraining a foundation model is always time-consumed due to the large scale of both the model parameter and training dataset. Besides being comput…

GPU