paper-with-me

Papers

CASSINI: Network-Aware Job Scheduling in Machine Learning Clusters

2023-08-01 · Sudarsanan Rajasekaran, Manya Ghobadi, Aditya Akella

We present CASSINI, a network-aware job scheduler for machine learning (ML) clusters. CASSINI introduces a novel geometric abstraction to consider the communication pattern of different jobs while placing them on network links. To do so, CASSINI uses an affinity graph that finds a series of time-shift values to adjust the communication phases of a subset of jobs, such that the communication patterns of jobs sharing the same network link are interleaved with each other. Experiments with 13 common ML models on a 24-server testbed demonstrate that compared to the state-of-the-art ML schedulers, CASSINI improves the average and tail completion time of jobs by up to 1.6x and 2.5x, respectively. Moreover, we show that CASSINI reduces the number of ECN marked packets in the cluster by up to 33x.

📄 PDF Abstract BibTeX arXiv:2308.00852

Code (0)

등록된 구현이 없습니다.

Tasks

Scheduling

Similar Papers 제목 키워드 기반

Contour Detection in Cassini ISS images based on Hierarchical Extreme Learning Machine and Dense Conditional Random Field

2019-08-22 · Xiqi Yang, Qingfeng Zhang, Zhan Li

In Cassini ISS (Imaging Science Subsystem) images, contour detection is often performed on disk-resolved object to accurately locate their center. Thus, the contour detection is a key problem. Traditional edge detection …

Contour DetectionEdge Detectionobject-detectionObject Detection

Resource Heterogeneity-Aware and Utilization-Enhanced Scheduling for Deep Learning Clusters

2025-03-13 · Abeda Sultana, Nabin Pakka, Fei Xu, Xu Yuan 외

Scheduling deep learning (DL) models to train on powerful clusters with accelerators like GPUs and TPUs, presently falls short, either lacking fine-grained heterogeneity awareness or leaving resources substantially under…

Scheduling

I/O Burst Prediction for HPC Clusters using Darshan Logs

2023-08-20 · Ehsan Saeedizade, Roya Taheri, Engin Arslan

Understanding cluster-wide I/O patterns of large-scale HPC clusters is essential to minimize the occurrence and impact of I/O interference. Yet, most previous work in this area focused on monitoring and predicting task a…

Scheduling

Network Contention-Aware Cluster Scheduling with Reinforcement Learning

2023-10-31 · Junyeol Ryu, Jeongyoon Eo

With continuous advances in deep learning, distributed training is becoming common in GPU clusters. Specifically, for emerging workloads with diverse amounts, ratios, and patterns of communication, we observe that networ…

GPUreinforcement-learningReinforcement LearningScheduling

Machine Learning Applications to Kronian Magnetospheric Reconnection Classification

2021-04-01 · Tadhg M. Garton, Caitriona M. Jackman, Andy W. Smith, Kiley L. Yeakel 외

The products of magnetic reconnection in Saturn's magnetotail are identified in magnetometer observations primarily through characteristic deviations in the north-south component of the magnetic field. These magnetic def…

BIG-bench Machine LearningClassificationGeneral Classification