paper-with-me

Papers

Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation

2026-05-08 · Joon Ha Kim, Geon-Woo Kim, Anoop Rachakonda, Daehyeok Kim arxiv

Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, since no single choice performs best across all workloads. Profile-based simulators are the standard tool, yet they hardcode their operation set to a specific configuration and re-profile every operation from scratch, making exploration prohibitively expensive. This cost stems from a missing structural understanding: every input dimension of each operation is fixed by the model configuration or determined by the incoming request. Many model-configuration values (e.g., head size, layer count) recur across models, so the same operation runs in many configurations; a single sweep over the request-dependent dimensions can serve them all. We present Dooly, which exploits this structure to achieve configuration-agnostic, redundancy-aware profiling. Dooly performs a single inference pass, labels each input dimension with its origin via taint propagation, and selectively profiles only operations absent from its latency database; stateful operations such as attention are isolated by reusing the serving engine's own initialization code, eliminating manual instrumentation. It builds latency regression models based on the database, which becomes a drop-in backend for existing simulators. Across two GPU platforms, three attention backends, and diverse model architectures, Dooly achieves simulation accuracy within 5% MAPE for TTFT and 8% for TPOT while reducing profiling GPU-hours by 56.4% across 12 models compared to the existing profiling approach. We have open-sourced Dooly at https://github.com/dooly-project.

📄 PDF Abstract BibTeX arXiv:2605.07985

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts

2026-04-28 · Vyom Sharma, Debajyoti Datta arxiv

The optimal kernel configuration for Mixture-of-Experts (MoE) inference depends on both batch size and the expert routing distribution, yet production systems dispatch from batch size alone, leaving 10-70% of kernel thro…

FleetSieve: Decision-Critical Profiling for SLO-Aware LLM Fleet Configuration

2026-08-20 · Huang Cheng, Scott Zhang, Aubert Li arxiv

Choosing tensor-parallel (TP) degrees and replica counts for an LLM serving fleet is difficult because performance is not monotonic in TP and the feasible choice can change with load. Exhaustive profiling resolves this u…

EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization

2026-05-14 · Zhiye Song, Kyungmi Lee, Eun Kyung Lee, Xin Zhang 외 arxiv

We present EnergyLens, an end-to-end framework for energy-aware large language model (LLM) inference optimization. As LLMs scale, predicting and reducing their energy footprint has become critical for sustainability and …

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

2026-07-28 · Wenjie Zhou, Yunting Liu, Renjiao Tang, Mark Wilson arxiv

Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too acc…

Karasu: A Collaborative Approach to Efficient Cluster Configuration for Big Data Analytics

2023-08-22 · Dominik Scheinert, Philipp Wiesner, Thorsten Wittkopp, Lauritz Thamsen 외

Selecting the right resources for big data analytics jobs is hard because of the wide variety of configuration options like machine type and cluster size. As poor choices can have a significant impact on resource efficie…