paper-with-me

홈 › Papers

NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference

2026-06-02 · Mubarak Adetunji Ojewale arxiv

Disaggregated LLM inference forces the KV cache to traverse the datacenter network before decoding begins, so transfer time enters directly into the Time to First Token (TTFT) budget. Current schedulers route on compute load and prefix-cache locality alone, ignoring the topological distance and dynamic congestion between prefill and decode instances. We close this gap with a thin operator-to-scheduler interface, the network cost oracle, and we prove that ignoring the network term renders cache-aware-only scheduling arbitrarily suboptimal as context length grows. NetKV, the O(|D|) per-request greedy that consumes this oracle, has tier rankings that are provably robust to stale telemetry. On a 64-GPU four-tier fat-tree simulator driven by Mooncake traces, NetKV reduces mean TTFT by up to 21.2% over round-robin and 17.6% over a tuned cache+load-aware scheduler, lifts SLO attainment by up to 20.1 percentage points, and keeps the Time Between Tokens overhead below 0.5 ms in every condition tested, with no changes to the transport, inference engine, or hardware.

📄 PDF Abstract BibTeX arXiv:2606.03910

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing

2025-09-05 · Jiahuan Yu, Aryan Taneja, Junfeng Lin, Minjia Zhang arxiv

The energy cost of Large Language Model (LLM) inference is rapidly becoming a barrier to sustainable and scalable deployment. Although modern serving architectures expose distinct prefill and decode behaviors, existing s…

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

2026-07-01 · Sangjin Choi, Sukmin Cho, Yifan Xiong, Ziyue Yang 외 hf

In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing decode routers balance only load; for mixture-of-experts (MoE) models this is incomplete: equally load…

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

2026-07-02 · Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj, Renee St. Amant 외 arxiv

Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates a new asymmetry: under bursty, heavy-tailed workloads prefill nodes saturate whil…

StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving

2026-02-11 · Satyam Kumar, Arpit Singh Gautam, Kailash Talreja, Saurabh Jha arxiv

Efficient LLM serving must balance throughput and latency across diverse, bursty workloads. We introduce StreamServe, a disaggregated prefill decode serving architecture that combines metric aware routing across compute …

FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling

2025-04-03 · Weiqing Li, Guochao Jiang, Xiangyong Ding, Zhangcheng Tao 외

Disaggregated inference has become an essential framework that separates the prefill (P) and decode (D) stages in large language model inference to improve throughput. However, the KV cache transfer faces significant del…

Language ModelingLanguage ModellingLarge Language ModelScheduling