paper-with-me

홈 › Papers

Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference

2026-03-04 · Patrick Wilhelm, Thorsten Wittkopp, Odej Kao arxiv

Large Language Models (LLMs) demonstrate exceptional performance across diverse tasks but come with substantial energy and computational costs, particularly in request-heavy scenarios. In many real-world applications, the full scale and capabilities of LLMs are often unnecessary, as Small Language Models (SLMs) can provide accurate responses for simpler text generation tasks. When enhanced with advanced reasoning strategies, such as Chain-of-Thought (CoT) prompting or Majority Voting, SLMs can approach the performance of larger models while reducing overall computational requirements. However, these strategies can also introduce additional energy costs, creating an energy-accuracy trade-off. Our analysis examines these trade-offs in test-time compute strategies for smaller models compared to larger ones, using the MMLU benchmark. Additionally, we explore the input-output token dynamics of transformer architectures, which result in nonlinear hardware energy operation curves for LLMs. To bridge AI research with its physical impact, we propose \textit{energy efficiency metrics}, including Energy-per-Token, as complements to traditional accuracy benchmarks. Beyond model selection, we propose controlled reasoning in CoT token generation, using operating curves to regulate reasoning depth dynamically. This vision integrates a energy-aware routing mechanism, ensuring that model selection and inference strategies balance accuracy for sustainable AI deployment.

📄 PDF Abstract BibTeX arXiv:2603.20224

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Membrane Interactions in Alzheimer`s Treatment Strategies with Multitarget Molecules

2024-03-10 · Pablo Zambrano

Addressing Alzheimer's disease (AD) requires innovative strategies beyond current single-target drugs. This Letter to the Editor suggests that multitarget molecules, especially those targeting neuronal membrane protectio…

Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems

2025-03-09 · Dany Moshkovich, Hadar Mulian, Sergey Zeltyn, Natti Eder 외

The rise of agentic AI systems, where agents collaborate to perform diverse tasks, poses new challenges with observing, analyzing and optimizing their behavior. Traditional evaluation and benchmarking approaches struggle…

Benchmarking

Order-book modelling and market making strategies

2018-06-13

Market making is one of the most important aspects of algorithmic trading, and it has been studied quite extensively from a theoretical point of view. The practical implementation of so-called "optimal strategies" howeve…

Algorithmic Trading

Advocating Feedback Control for Human-Earth System Applications

2024-05-12 · Guido Cavraro

This paper proposes a feedback control perspective for Human-Earth Systems (HESs) which essentially are complex systems that capture the interactions between humans and nature. Recent attention in HES research has been d…

Rethinking Diversity in Deep Neural Network Testing

2023-05-25 · Zi Wang, Jihye Choi, Ke Wang, Somesh Jha

Motivated by the success of traditional software testing, numerous diversity measures have been proposed for testing deep neural networks (DNNs). In this study, we propose a shift in perspective, advocating for the consi…

DiversityDNN Testingsoftware testing