paper-with-me

홈 › Papers

GOGH: Correlation-Guided Orchestration of GPUs in Heterogeneous Clusters

2025-10-17 · Ahmad Raeisi, Mahdi Dolati, Sina Darabi, Sadegh Talebi, Patrick Eugster, Ahmad Khonsari arxiv

The growing demand for computational resources in machine learning has made efficient resource allocation a critical challenge, especially in heterogeneous hardware clusters where devices vary in capability, age, and energy efficiency. Upgrading to the latest hardware is often infeasible, making sustainable use of existing, mixed-generation resources essential. In this paper, we propose a learning-based architecture for managing machine learning workloads in heterogeneous clusters. The system operates online, allocating resources to incoming training or inference requests while minimizing energy consumption and meeting performance requirements. It uses two neural networks: the first provides initial estimates of how well a new model will utilize different hardware types and how it will affect co-located models. An optimizer then allocates resources based on these estimates. After deployment, the system monitors real performance and uses this data to refine its predictions via a second neural network. This updated model improves estimates not only for the current hardware but also for hardware not initially allocated and for co-location scenarios not yet observed. The result is an adaptive, iterative approach that learns over time to make more effective resource allocation decisions in heterogeneous deep learning clusters.

📄 PDF Abstract BibTeX arXiv:2510.15652

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization

2025-01-16 · Zixun Fang, Zhiheng Liu, Kai Zhu, Yu Liu 외

Video colorization aims to transform grayscale videos into vivid color representations while maintaining temporal consistency and structural integrity. Existing video colorization methods often suffer from color bleeding…

ColorizationOptical Flow Estimation

HALO 1.0: A Hardware-agnostic Accelerator Orchestration Framework for Enabling Hardware-agnostic Programming with True Performance Portability for Heterogeneous HPC

2020-11-22 · Michael Riera, Erfan Bank Tavakoli, Masudul Hassan Quraishi, Fengbo Ren

This paper presents HALO 1.0, an open-ended extensible multi-agent software framework that implements a set of proposed hardware-agnostic accelerator orchestration (HALO) principles. HALO implements a novel compute-centr…

Optimal Kernel Orchestration for Tensor Programs with Korch

2024-06-13 · Muyan Hu, Ashwin Venkatram, Shreyashri Biswas, Balamurugan Marimuthu 외

Kernel orchestration is the task of mapping the computation defined in different operators of a deep neural network (DNN) to the execution of GPU kernels on modern hardware platforms. Prior approaches optimize kernel orc…

DiversityGPUtensor algebra

ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads

2026-04-07 · Jingwei Zuo, Xinze Feng, Zien Liu, Kaijian Wang 외 arxiv

Low-Rank Adaptation (LoRA) is now the dominant method for parameter-efficient fine-tuning of large language models, but achieving a high-quality adapter often requires systematic hyperparameter tuning because LoRA perfor…

parameter-efficient fine-tuning

AIvailable: A Software-Defined Architecture for LLM-as-a-Service on Heterogeneous and Legacy GPUs

2025-11-06 · Pedro Antunes, Ana Rita Ortigoso, Gabriel Vieira, Daniel Fuentes 외 arxiv

The rise of Large Language Models (LLM) has increased the need for scalable, high-performance inference systems, yet most existing frameworks assume homogeneous, resource-rich hardware, often unrealistic in academic, or …