paper-with-me

홈 › Papers

Energy-Aware Routing to Large Reasoning Models

2025-12-23 · Austin R. Ellis-Mohr, Max Hartman, Lav R. Varshney arxiv

Large reasoning models (LRMs) have heterogeneous inference energy costs based on which model is used and how much it reasons. To reduce energy, it is important to choose the right LRM and operate it in the right way. As a result, the performance of systems that dispatch tasks to different individual LRMs depend on the balance between mean energy provisioning and stochastic fluctuations. The critical regime is the unique operating point at which neither auxiliary energy nor baseline energy is systematically wasted. Increasing baseline supply shifts the system toward persistent over-supply and baseline-energy waste, while reducing supply induces persistent reliance on auxiliary energy. Yet in this regime, performance remains volatility-limited and so a second-order characterization provides further insights that we develop. Here, performance is governed by how variability is absorbed across time, models, and execution choices. This perspective highlights variance-aware routing and dispatch as a principled design axis, and provides a theoretical basis for developing energy-aware model routing policies. Routing behavior is characterized when dispatch policies are based on training-compute and inference-compute scaling laws for LRMs.

📄 PDF Abstract BibTeX arXiv:2601.00823

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

StructuredDNA: A Bio-Physical Framework for Energy-Aware Transformer Routing

2025-12-01 · Mustapha Hamdi arxiv

The rapid scaling of large computational models has led to a critical increase in energy and compute costs. Inspired by biological systems where structure and function emerge from low-energy configurations, we introduce …

INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference

2026-05-13 · Ahmed Šabanović, Paul Joe Maliakel, Ivona Brandić arxiv

Edge deployment of Vision-Language Models (VLMs) faces a tradeoff between latency and accuracy: cloud execution provides high-quality predictions but incurs communication delay and energy cost, while edge-only execution …

Visual Question Answering

Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference

2026-03-04 · Patrick Wilhelm, Thorsten Wittkopp, Odej Kao arxiv

Large Language Models (LLMs) demonstrate exceptional performance across diverse tasks but come with substantial energy and computational costs, particularly in request-heavy scenarios. In many real-world applications, th…

Text Generation

Robust Energy-Aware Routing for Air-Ground Cooperative Multi-UAV Delivery in Wind-Uncertain Environments

2026-04-15 · Tianshun Li, Hongliang Lu, Yanggang Sheng, Zhongzhen Wang 외 arxiv

Ensuring energy feasibility under wind uncertainty is critical for the safety and reliability of UAV delivery missions. In realistic truck-drone logistics systems, UAVs must deliver parcels and safely return under time-v…

ECVL-ROUTER: Scenario-Aware Routing for Vision-Language Models

2025-10-31 · Xin Tang, Youfang Han, Fangfei Gou, Wei Zhao 외 arxiv

Vision-Language Models (VLMs) excel in diverse multimodal tasks. However, user requirements vary across scenarios, which can be categorized into fast response, high-quality output, and low energy consumption. Relying sol…