paper-with-me

홈 › Papers

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations

2026-07-10 · Nada Zine, Tristan Coignion, Vincenzo Stoico, Clément Quinton, Romain Rouvoy, Patricia Lago arxiv

Large Language Models are reshaping how software is developed and maintained. They are typically deployed in production using inference engines such as vLLM, which can efficiently serve pre-trained, highly configurable models. While prior work has focused on model architectures and hardware acceleration, the impact of inference engine configuration on energy consumption, performance, and output quality remains poorly understood. In this paper, we present a large-scale controlled study of three selected vLLM configuration options: attention kernel type, prefix caching, and chunked prefill. We evaluate all combinations of these configurations across 5 open-weight LLMs and 5 diverse inference tasks, totaling $9,000$ runs and $93,600$ measures. We analyze energy consumption, latency, and accuracy, and examine both main effects and interaction effects between configuration options and tasks. Our results show that the studied configuration options significantly impact energy and performance, mainly driven by attention type and prefix caching, while chunked prefill has a limited effect under the default vLLM serving configuration and evaluated workloads. These effects are highly model- and workload-dependent, and no configuration is universally optimal. We further show that model choice dominates global trade-offs, while configuration tuning provides local improvements along the Pareto frontier. Unexpectedly, inference options can also affect model accuracy.

📄 PDF Abstract BibTeX arXiv:2607.09172

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Watt For What: Rethinking Deep Learning's Energy-Performance Relationship

2023-10-10 · Shreyank N Gowda, Xinyue Hao, Gen Li, Shashank Narayana Gowda 외

Deep learning models have revolutionized various fields, from image recognition to natural language processing, by achieving unprecedented levels of accuracy. However, their increasing energy consumption has raised conce…

Deep Learning

Hardware-Based Microgrid Coupled to Real-Time Simulated Power Grids for Evaluating New Control Strategies in Future Energy Systems

2024-09-03 · Michael Kyesswa, Friedrich Wiegel, Jan Wachter, Uwe Kühnapfel 외

The design of new control strategies for future energy systems can neither be directly tested in real power grids nor be evaluated based on only current grid situations. In this regard, extensive tests are required in la…

Hydrogen production from blended waste biomass: pyrolysis, thermodynamic-kinetic analysis and AI-based modelling

2025-10-11 · Sana Kordoghli, Abdelhakim Settar, Oumayma Belaati, Mohammad Alkhatib 외 arxiv

This work contributes to advancing sustainable energy and waste management strategies by investigating the thermochemical conversion of food-based biomass through pyrolysis, highlighting the role of artificial intelligen…

Evaluating Different Fault Injection Abstractions on the Assessment of DNN SW Hardening Strategies

2024-12-11 · Giuseppe Esposito, Juan David Guerrero-Balaguera, Josie Esteban Rodriguez Condia, Matteo Sonza Reorda

The reliability of Neural Networks has gained significant attention, prompting efforts to develop SW-based hardening techniques for safety-critical scenarios. However, evaluating hardening techniques using application-le…

EECD-Net: Energy-Efficient Crack Detection with Spiking Neural Networks and Gated Attention

2025-06-05 · Shuo Zhang

Crack detection on road surfaces is a critical measurement technology in the instrumentation domain, essential for ensuring infrastructure safety and transportation reliability. However, due to limited energy and low-res…

Super-Resolution