paper-with-me

Papers

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

2026-07-02 · Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj, Renee St. Amant, Victor Rühle arxiv

Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates a new asymmetry: under bursty, heavy-tailed workloads prefill nodes saturate while decode nodes have compute underutilized, and on a production-style A100 cluster with 2 prefill and 2 decode nodes (2P2D), we find that prefill execution accounts for only 2-23% of P95 Time-to-First-Token (TTFT). Queuing and inter-node GPU-GPU KV-cache transfer account for the rest. We present a proactive prefill-deflecting scheduler that lets decode nodes serve prefill phase of requests as chunked-prefill steps interleaved with their in-flight decode batches. For each queued request, we estimate the TTFT it would see on the prefill node, and on every decode node, search for the largest chunk schedule that keeps in-flight decodes within their Time-Between-Tokens (TBT) SLO and deflect when the decode path helps tail latency. Because the prefill phase of deflected requests runs in place on the decode node, the inter-node KV transfer is eliminated. Implemented on vLLM and evaluated on production-style traces with DeepSeek-V2-Lite, our approach reduces P95 TTFT by upto 81% and raises SLO attainment by upto 79% over state-of-the-art disaggregated schedulers, at sub-millisecond per-request routing cost.

📄 PDF Abstract BibTeX arXiv:2607.02043

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

2026-07-01 · Sangjin Choi, Sukmin Cho, Yifan Xiong, Ziyue Yang 외 hf

In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing decode routers balance only load; for mixture-of-experts (MoE) models this is incomplete: equally load…

PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving

2026-02-12 · Sunghyeon Woo, Hoseung Kim, Sunghwan Shim, Minjung Jo 외 arxiv

Multi-agent systems increasingly orchestrate multiple specialized language models to solve complex real-world problems, often invoking them over a shared context. This execution pattern repeatedly processes the same prom…

Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications

2025-11-14 · Jiaxi Li, Yue Zhu, Eun Kyung Lee, Klara Nahrstedt arxiv

Different from traditional Large Language Model (LLM) serving that colocates the prefill and decode stages on the same GPU, disaggregated serving dedicates distinct GPUs to prefill and decode workload. Once the prefill G…

StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving

2026-02-11 · Satyam Kumar, Arpit Singh Gautam, Kailash Talreja, Saurabh Jha arxiv

Efficient LLM serving must balance throughput and latency across diverse, bursty workloads. We introduce StreamServe, a disaggregated prefill decode serving architecture that combines metric aware routing across compute …

VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing

2025-09-05 · Jiahuan Yu, Aryan Taneja, Junfeng Lin, Minjia Zhang arxiv

The energy cost of Large Language Model (LLM) inference is rapidly becoming a barrier to sustainable and scalable deployment. Although modern serving architectures expose distinct prefill and decode behaviors, existing s…