paper-with-me

홈 › Papers

Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale

2024-02-25 · Dan Zhao, Siddharth Samsi, Joseph McDonald, Baolin Li, David Bestor, Michael Jones, Devesh Tiwari, Vijay Gadepally

As research and deployment of AI grows, the computational burden to support and sustain its progress inevitably does too. To train or fine-tune state-of-the-art models in NLP, computer vision, etc., some form of AI hardware acceleration is virtually a requirement. Recent large language models require considerable resources to train and deploy, resulting in significant energy usage, potential carbon emissions, and massive demand for GPUs and other hardware accelerators. However, this surge carries large implications for energy sustainability at the HPC/datacenter level. In this paper, we study the aggregate effect of power-capping GPUs on GPU temperature and power draw at a research supercomputing center. With the right amount of power-capping, we show significant decreases in both temperature and power draw, reducing power consumption and potentially improving hardware life-span with minimal impact on job performance. While power-capping reduces power draw by design, the aggregate system-wide effect on overall energy consumption is less clear; for instance, if users notice job performance degradation from GPU power-caps, they may request additional GPU-jobs to compensate, negating any energy savings or even worsening energy consumption. To our knowledge, our work is the first to conduct and make available a detailed analysis of the effects of GPU power-capping at the supercomputing scale. We hope our work will inspire HPCs/datacenters to further explore, evaluate, and communicate the impact of power-capping AI hardware accelerators for more sustainable AI.

📄 PDF Abstract BibTeX arXiv:2402.18593

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Similar Papers 제목 키워드 기반

A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale

2024-10-07 · Wesley Brewer, Matthias Maiterth, Vineet Kumar, Rafal Wojda 외

We present ExaDigiT, an open-source framework for developing comprehensive digital twins of liquid-cooled supercomputers. It integrates three main modules: (1) a resource allocator and power simulator, (2) a transient th…

The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures

2026-05-12 · Bole Ma, Ayesha Afzal, Jan Eitzinger, Gerhard Wellein arxiv

Power capping is the standard GPU energy lever in LLM serving, and it appears to work: throughput drops, power readings fall, and energy budgets are met. We show the appearance is illusory for the phase that dominates pr…

Great Power, Great Responsibility: Recommendations for Reducing Energy for Training Language Models

2022-05-19 · Findings (NAACL) 2022 7 · Joseph McDonald, Baolin Li, Nathan Frey, Devesh Tiwari 외

The energy requirements of current natural language processing models continue to grow at a rapid, unsustainable pace. Recent works highlighting this problem conclude there is an urgent need for methods that reduce the e…

Cloud ComputingGPULanguage ModelingLanguage Modelling

Reducing the Barriers to Entry for Foundation Model Training

2024-04-12 · Paolo Faraboschi, Ellis Giles, Justin Hotard, Konstanty Owczarek 외

The world has recently witnessed an unprecedented acceleration in demands for Machine Learning and Artificial Intelligence applications. This spike in demand has imposed tremendous strain on the underlying technology sta…

GPU

JUWELS Booster -- A Supercomputer for Large-Scale AI Research

2021-06-30 · Stefan Kesselheim, Andreas Herten, Kai Krajsek, Jan Ebert 외

In this article, we present JUWELS Booster, a recently commissioned high-performance computing system at the J\"ulich Supercomputing Center. With its system architecture, most importantly its large number of powerful Gra…