Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
With the ubiquitous use of modern large language models (LLMs) across industries, the inference serving for these models is ever expanding. Given the high compute and memory requirements of modern LLMs, more and more top-of-the-line GPUs are being deployed to serve these models. Energy availability has come to the forefront as the biggest challenge for data center expansion to serve these models. In this paper, we present the trade-offs brought up by making energy efficiency the primary goal of LLM serving under performance SLOs. We show that depending on the inputs, the model, and the service-level agreements, there are several knobs available to the LLM inference provider to use for being energy efficient. We characterize the impact of these knobs on the latency, throughput, as well as the energy. By exploring these trade-offs, we offer valuable insights into optimizing energy usage without compromising on performance, thereby paving the way for sustainable and cost-effective LLM deployment in data center environments.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Ethics Whitepaper: Whitepaper on Ethical Research into Large Language Models
This whitepaper offers an overview of the ethical considerations surrounding research into or with large language models (LLMs). As LLMs become more integrated into widely used applications, their societal impact increas…
EthicsUnderstanding Efficiency: Quantization, Batching, and Serving Strategies in LLM Energy Use
Large Language Models (LLMs) are increasingly deployed in production, contributing towards shifting the burden in terms of computational resources and energy demands from training to inference. While prior work has exami…
Text GenerationEnergy efficiency of DMAs vs. conventional MIMO: a sensitivity analysis
Motivated by the stringent and challenging need for `greener communications' in increasingly power-hungry 5G networks, this paper presents a detailed energy efficiency analysis for three different multi-antenna architect…
SensitivityImpact of ML Optimization Tactics on Greener Pre-Trained ML Models
Background: Given the fast-paced nature of today's technology, which has surpassed human performance in tasks like image classification, visual reasoning, and English understanding, assessing the impact of Machine Learni…
GPUimage-classificationImage ClassificationQuantization+1GA4GC: Greener Agent for Greener Code via Multi-Objective Configuration Optimization
Coding agents powered by LLMs face critical sustainability and scalability challenges in industrial deployment, with single runs consuming over 100k tokens and incurring environmental costs that may exceed optimization b…