paper-with-me

홈 › Papers

Dynamic Scheduling of MPI-based Distributed Deep Learning Training Jobs

2019-08-21 · Tim Capes, Vishal Raheja, Mete Kemertas, Iqbal Mohomed

There is a general trend towards solving problems suited to deep learning with more complex deep learning architectures trained on larger training sets. This requires longer compute times and greater data parallelization or model parallelization. Both data and model parallelism have been historically faster in parameter server architectures, but data parallelism is starting to be faster in ring architectures due to algorithmic improvements. In this paper, we analyze the math behind ring architectures and make an informed adaptation of dynamic scheduling to ring architectures. To do so, we formulate a non-convex, non-linear, NP-hard integer programming problem and a new efficient doubling heuristic for its solution. We build upon Horovod: an open source ring architecture framework over TensorFlow. We show that Horovod jobs have a low cost to stop and restart and that stopping and restarting ring architecture jobs leads to faster completion times. These two facts make dynamic scheduling of ring architecture jobs feasible. Lastly, we simulate a scheduler using these runs and show a more than halving of average job time on some workload patterns.

📄 PDF Abstract BibTeX arXiv:1908.08082

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningMathScheduling

Similar Papers 제목 키워드 기반

A Deep Reinforcement Learning Approach to Multi-component Job Scheduling in Edge Computing

2019-08-26 · Zhi Cao, Honggang Zhang, Yu Cao, Benyuan Liu

We are interested in the optimal scheduling of a collection of multi-component application jobs in an edge computing system that consists of geo-distributed edge computing nodes connected through a wide area network. The…

Deep Reinforcement LearningEdge-computingreinforcement-learningReinforcement Learning+2

Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters

2025-01-09 · Ziyue Luo, Jia Liu, Myungjin Lee, Ness B. Shroff

The recent explosive growth of deep learning (DL) models has necessitated a compelling need for efficient job scheduling for distributed deep learning training with mixed parallelisms (DDLwMP) in GPU clusters. This paper…

GPUScheduling

Communication Contention Aware Scheduling of Multiple Deep Learning Training Jobs

2020-02-24 · Qiang Wang, Shaohuai Shi, Canhui Wang, Xiaowen Chu

Distributed Deep Learning (DDL) has rapidly grown its popularity since it helps boost the training performance on high-performance GPU clusters. Efficient job scheduling is indispensable to maximize the overall performan…

Deep LearningGPUScheduling

Venn: Resource Management for Collaborative Learning Jobs

2023-12-13 · Jiachen Liu, Fan Lai, Ding Ding, Yiwen Zhang 외

In recent years, collaborative learning (CL) has emerged as a promising approach for machine learning (ML) and data science across distributed edge devices. As the deployment of CL jobs increases, they inevitably contend…

Federated LearningManagementScheduling

Reducing Fragmentation and Starvation in GPU Clusters through Dynamic Multi-Objective Scheduling

2025-12-04 · Akhmadillo Mamirov arxiv

GPU clusters have become essential for training and deploying modern AI systems, yet real deployments continue to report average utilization near 50%. This inefficiency is largely caused by fragmentation, heterogeneous w…