paper-with-me

홈 › Papers

HELIOS: Adaptive Model And Early-Exit Selection for Efficient LLM Inference Serving

2025-04-14 · Avinash Kumar, Shashank Nag, Jason Clemons, Lizy John, Poulami Das

Deploying large language models (LLMs) presents critical challenges due to the inherent trade-offs associated with key performance metrics, such as latency, accuracy, and throughput. Typically, gains in one metric is accompanied with degradation in others. Early-Exit LLMs (EE-LLMs) efficiently navigate this trade-off space by skipping some of the later model layers when it confidently finds an output token early, thus reducing latency without impacting accuracy. However, as the early exits taken depend on the task and are unknown apriori to request processing, EE-LLMs conservatively load the entire model, limiting resource savings and throughput. Also, current frameworks statically select a model for a user task, limiting our ability to adapt to changing nature of the input queries. We propose HELIOS to address these challenges. First, HELIOS shortlists a set of candidate LLMs, evaluates them using a subset of prompts, gathering telemetry data in real-time. Second, HELIOS uses the early exit data from these evaluations to greedily load the selected model only up to a limited number of layers. This approach yields memory savings which enables us to process more requests at the same time, thereby improving throughput. Third, HELIOS monitors and periodically reassesses the performance of the candidate LLMs and if needed, switches to another model that can service incoming queries more efficiently (such as using fewer layers without lowering accuracy). Our evaluations show that HELIOS achieves 1.48$\times$ throughput, 1.10$\times$ energy-efficiency, 1.39$\times$ lower response time, and 3.7$\times$ improvements in inference batch sizes compared to the baseline, when optimizing for the respective service level objectives.

📄 PDF Abstract BibTeX arXiv:2504.10724

Code (0)

등록된 구현이 없습니다.

Tasks

Navigate

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

HELIOS: Harmonizing Early Fusion, Late Fusion, and LLM Reasoning for Multi-Granular Table-Text Retrieval

2026-02-25 · Sungho Park, Joohyung Yun, Jongwuk Lee, Wook-Shin Han arxiv

Table-text retrieval aims to retrieve relevant tables and text to support open-domain question answering. Existing studies use either early or late fusion, but face limitations. Early fusion pre-aligns a table row with i…

Open-Domain Question AnsweringText Retrieval

Adaptive Deep Neural Network Inference Optimization with EENet

2023-01-15 · Fatih Ilhan, Ka-Ho Chow, Sihao Hu, Tiansheng Huang 외

Well-trained deep neural networks (DNNs) treat all test samples equally during prediction. Adaptive DNN inference with early exiting leverages the observation that some test examples can be easier to predict than others.…

Inference OptimizationSchedulingSST-2

Early-stopped aggregation: Adaptive inference with computational efficiency

2026-04-15 · Ilsang Ohn, Shitao Fan, Jungbin Jun, Lizhen Lin arxiv

When considering a model selection or, more generally, an aggregation approach for adaptive statistical inference, it is often necessary to compute estimators over a wide range of model complexities including unnecessari…

Computational Efficiency

Virtual laser scanning with HELIOS++: A novel take on ray tracing-based simulation of topographic 3D laser scanning

2021-01-21 · Lukas Winiwarter, Alberto Manuel Esmorís Pena, Hannah Weiser, Katharina Anders 외

Topographic laser scanning is a remote sensing method to create detailed 3D point cloud representations of the Earth's surface. Since data acquisition is expensive, simulations can complement real data given certain prem…

When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth

2026-04-17 · Dongxin Guo, Jikun Wu, Siu Ming Yiu arxiv

Early-exit neural networks enable adaptive computation by allowing confident predictions to exit at intermediate layers, achieving 2-8$\times$ inference speedup. Despite widespread deployment, their generalization proper…