paper-with-me

홈 › Papers

Semantic-Aware Scheduling for GPU Clusters with Large Language Models

2025-10-02 · Zerui Wang, Qinghao Hu, Ana Klimovic, Tianwei Zhang, Yonggang Wen, Peng Sun, Dahua Lin arxiv

Deep learning (DL) schedulers are pivotal in optimizing resource allocation in GPU clusters, but operate with a critical limitation: they are largely blind to the semantic context of the jobs they manage. This forces them to rely on limited metadata, leading to high profiling overhead, unreliable duration estimation, inadequate failure handling, and poor observability. To this end, we propose SchedMate, a framework that bridges this semantic gap by systematically extracting deep insights from overlooked, unstructured data sources: source code, runtime logs, and historical jobs. SchedMate enhances existing schedulers non-intrusively through three LLM-based components. Our implementation integrates seamlessly with existing deep learning schedulers. Evaluations on a 128-GPU physical cluster and extensive simulations on production traces show SchedMate reduces average job completion times by up to 1.91x, substantially enhancing the scheduling performance, demonstrating the critical role of semantic-awareness in modern DL scheduling.

📄 PDF Abstract BibTeX arXiv:2510.03334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters

2025-09-25 · Lingling Zeng, Gen Zhang, Jialin Peng, Xiang Xu 외 arxiv

As AI cluster sizes continue to expand and the demand for large-language-model (LLM) training and inference workloads grows rapidly, traditional scheduling systems face significant challenges in balancing resource utiliz…

Semantic Scheduling for LLM Inference

2025-06-13 · Wenyue Hua, Dujian Ding, Yile Gu, Yujie Ren 외

Conventional operating system scheduling algorithms are largely content-ignorant, making decisions based on factors such as latency or fairness without considering the actual intents or semantics of processes. Consequent…

FairnessManagementScheduling

Network Contention-Aware Cluster Scheduling with Reinforcement Learning

2023-10-31 · Junyeol Ryu, Jeongyoon Eo

With continuous advances in deep learning, distributed training is becoming common in GPU clusters. Specifically, for emerging workloads with diverse amounts, ratios, and patterns of communication, we observe that networ…

GPUreinforcement-learningReinforcement LearningScheduling

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning

2026-06-21 · Zicheng Xu, Ruixuan Zhang, Yu-Neng Chuang, Xiuyi Lou 외 arxiv

Large Language Models (LLMs) achieve remarkable reasoning capabilities through reinforcement learning (RL) post-training. However, existing RL post-training commonly relies on uniform data sampling, which ignores the sem…

Reinforcement Learning

Resource Heterogeneity-Aware and Utilization-Enhanced Scheduling for Deep Learning Clusters

2025-03-13 · Abeda Sultana, Nabin Pakka, Fei Xu, Xu Yuan 외

Scheduling deep learning (DL) models to train on powerful clusters with accelerators like GPUs and TPUs, presently falls short, either lacking fine-grained heterogeneity awareness or leaving resources substantially under…

Scheduling