paper-with-me

홈 › Papers

PARIS and ELSA: An Elastic Scheduling Algorithm for Reconfigurable Multi-GPU Inference Servers

2022-02-27 · Yunseong Kim, Yujeong Choi, Minsoo Rhu

In cloud machine learning (ML) inference systems, providing low latency to end-users is of utmost importance. However, maximizing server utilization and system throughput is also crucial for ML service providers as it helps lower the total-cost-of-ownership. GPUs have oftentimes been criticized for ML inference usages as its massive compute and memory throughput is hard to be fully utilized under low-batch inference scenarios. To address such limitation, NVIDIA's recently announced Ampere GPU architecture provides features to "reconfigure" one large, monolithic GPU into multiple smaller "GPU partitions". Such feature provides cloud ML service providers the ability to utilize the reconfigurable GPU not only for large-batch training but also for small-batch inference with the potential to achieve high resource utilization. In this paper, we study this emerging GPU architecture with reconfigurability to develop a high-performance multi-GPU ML inference server. Our first proposition is a sophisticated partitioning algorithm for reconfigurable GPUs that systematically determines a heterogeneous set of multi-granular GPU partitions, best suited for the inference server's deployment. Furthermore, we co-design an elastic scheduling algorithm tailored for our heterogeneously partitioned GPU server which effectively balances low latency and high GPU utilization.

📄 PDF Abstract BibTeX arXiv:2202.13481

Code (0)

등록된 구현이 없습니다.

Tasks

GPUScheduling

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing

2026-05-20 · Kang You, Chen Nie, Lee Jun Yan, Ziling Wei 외 arxiv

Spiking neural networks (SNNs) exploit event-driven and addition-only computation to substantially improve efficiency for intelligent computation. A key temporal property of SNNs, elastic inference, allows outputs to eme…

ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

2026-07-07 · Tianjiao Yu, Xinzhuo Li, Yifan Shen, Onkar Susladkar 외 arxiv

Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D interaction remains largely implicit. Existing methods concatenate text and 3D token…

3D Generation

ELSA: A Throughput-Optimized Design of an LSTM Accelerator for Energy-Constrained Devices

2019-10-19 · Elham Azari, Sarma Vrudhula

The next significant step in the evolution and proliferation of artificial intelligence technology will be the integration of neural network (NN) models within embedded and mobile systems. This calls for the design of co…

Language ModelingLanguage Modelling

Predicting Risk of Dementia with Survival Machine Learning and Statistical Methods: Results on the English Longitudinal Study of Ageing Cohort

2023-06-17 · Daniel Stamate, Henry Musto, Olesya Ajnakina, Daniel Stahl

Machine learning models that aim to predict dementia onset usually follow the classification methodology ignoring the time until an event happens. This study presents an alternative, using survival analysis within the co…

Survival Analysis

A Flexible Job Shop Scheduling Problem Involving Reconfigurable Machine Tools Under Industry 5.0

2024-10-16 · Hessam Bakhshi-Khaniki, Reza Tavakkoli-Moghaddam, Zdenek Hanzalek, Behdin Vahedi-Nouri

The rise of Industry 5.0 has introduced new demands for manufacturing companies, requiring a shift in how production schedules are managed to address human centered, environmental, and economic goals comprehensively. The…

Job Shop SchedulingScheduling