paper-with-me

Papers

DRAS-CQSim: A Reinforcement Learning based Framework for HPC Cluster Scheduling

2021-05-16 · Yuping Fan, Zhiling Lan

For decades, system administrators have been striving to design and tune cluster scheduling policies to improve the performance of high performance computing (HPC) systems. However, the increasingly complex HPC systems combined with highly diverse workloads make such manual process challenging, time-consuming, and error-prone. We present a reinforcement learning based HPC scheduling framework named DRAS-CQSim to automatically learn optimal scheduling policy. DRAS-CQSim encapsulates simulation environments, agents, hyperparameter tuning options, and different reinforcement learning algorithms, which allows the system administrators to quickly obtain customized scheduling policies.

📄 PDF Abstract BibTeX arXiv:2105.07526

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement LearningReinforcement Learning (RL)Scheduling

Similar Papers 제목 키워드 기반

Deep Reinforcement Agent for Scheduling in HPC

2021-02-11 · Yuping Fan, Zhiling Lan, Taylor Childers, Paul Rich 외

Cluster scheduler is crucial in high-performance computing (HPC). It determines when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human ex…

Deep Reinforcement LearningScheduling

Network Contention-Aware Cluster Scheduling with Reinforcement Learning

2023-10-31 · Junyeol Ryu, Jeongyoon Eo

With continuous advances in deep learning, distributed training is becoming common in GPU clusters. Specifically, for emerging workloads with diverse amounts, ratios, and patterns of communication, we observe that networ…

GPUreinforcement-learningReinforcement LearningScheduling

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning

2026-06-21 · Zicheng Xu, Ruixuan Zhang, Yu-Neng Chuang, Xiuyi Lou 외 arxiv

Large Language Models (LLMs) achieve remarkable reasoning capabilities through reinforcement learning (RL) post-training. However, existing RL post-training commonly relies on uniform data sampling, which ignores the sem…

Reinforcement Learning

Energy-aware Scheduling of Jobs in Heterogeneous Cluster Systems Using Deep Reinforcement Learning

2019-12-11 · Amirhossein Esmaili, Massoud Pedram

Energy consumption is one of the most critical concerns in designing computing devices, ranging from portable embedded systems to computer cluster systems. Furthermore, in the past decade, cluster systems have increasing…

Deep Reinforcement LearningManagementReinforcement LearningReinforcement Learning (RL)+1

Scalable Reinforcement Learning for Virtual Machine Scheduling

2025-03-01 · Junjie Sheng, Jiehao Wu, Haochuan Cui, Yiqiu Hu 외

Recent advancements in reinforcement learning (RL) have shown promise for optimizing virtual machine scheduling (VMS) in small-scale clusters. The utilization of RL to large-scale cloud computing scenarios remains notabl…

Cloud Computingreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1