paper-with-me

홈 › Papers

DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing

2025-11-06 · Lei Gao, Chaoyi Jiang, Hossein Entezari Zarch, Daniel Wong, Mark Hill, Murali Annavaram arxiv

Modern LLM serving systems must sustain high throughput while meeting strict latency SLOs across two distinct inference phases: compute-intensive prefill and memory-bound decode phases. Existing approaches either (1) aggregate both phases on shared GPUs, leading to interference between prefill and decode phases, which degrades Time-Between-Tokens (TBT); or (2) disaggregate the two phases across GPUs, improving latency but wasting resources through duplicated models and KV cache transfers. We present DuetServe, a unified LLM serving framework that achieves disaggregation-level isolation within a single GPU. DuetServe operates in aggregated mode by default and dynamically activates SM-level GPU spatial multiplexing when TBT degradation is predicted. Its key idea is to decouple prefill and decode execution only when needed through fine-grained, adaptive SM partitioning that provides phase isolation only when contention threatens latency service level objectives. DuetServe integrates (1) an attention-aware roofline model to forecast iteration latency, (2) a partitioning optimizer that selects the optimal SM split to maximize throughput under TBT constraints, and (3) an interruption-free execution engine that eliminates CPU-GPU synchronization overhead. Evaluations show that DuetServe improves total throughput by up to 1.3x while maintaining low generation latency compared to state-of-the-art frameworks.

📄 PDF Abstract BibTeX arXiv:2511.04791

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sangam: Efficiently Serving Diffusion LLMs with the AR Stack

2026-07-05 · Nitin Kedia, Saurabh Agarwal, Myungjin Lee, Aditya Akella arxiv

Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation. Their bidirectional attention prevents exact autoregressive-style …

LAPS: A Length-Aware-Prefill LLM Serving System

2026-01-04 · Jianshu She, Zonghang Li, Hongchao Du, Shangyu Wu 외 arxiv

LAPS identifies and disaggregates requests with different prompt lengths in LLM serving to reduce TTFT latency. While recent systems have decoupled the prefill and decode stages to improve throughput, they still rely on …

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

2026-07-02 · Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj, Renee St. Amant 외 arxiv

Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates a new asymmetry: under bursty, heavy-tailed workloads prefill nodes saturate whil…

PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving

2026-02-12 · Sunghyeon Woo, Hoseung Kim, Sunghwan Shim, Minjung Jo 외 arxiv

Multi-agent systems increasingly orchestrate multiple specialized language models to solve complex real-world problems, often invoking them over a shared context. This execution pattern repeatedly processes the same prom…

RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation

2026-01-16 · Amna Masood, Pratishtha Gaur, Nuwan Jayasena arxiv

Two widely adopted techniques for LLM inference serving systems today are hybrid batching and disaggregated serving. A hybrid batch combines prefill and decode tokens of different requests in the same batch to improve re…