paper-with-me

홈 › Papers

PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving

2026-05-10 · Bing Xie, Zhipeng Wang, Masahiro Tanaka, Zheng Zhen arxiv

We present PEEK, a lightweight scheduling and eviction framework for both online (streaming) and offline (batch) LLM serving; this paper focuses on the online regime. PEEK maintains an incremental radix tree over the pending queue, exposing prefix-sharing clusters no existing engine surfaces. A low-overhead dual-walk matches the tree against the engine's prefix cache to yield longest-prefix-match for every waiting request; PEEK then admits cluster pioneers first so siblings inherit the freshly cached prefix, a co-designed eviction hook protects blocks ancestral to queued demand, and a multi-lane stride scheduler bounds starvation. On SGLang and vLLM across five workloads up to 4$\times$H100 (DP=2 over TP=2), PEEK delivers up to 3.0$\times$/2.6$\times$ cache hit, 7.9$\times$/7.1$\times$ TTFT, 6.7$\times$/5.5$\times$ E2E, and 3.6$\times$/4.5$\times$ throughput gains over each engine's strongest stock baseline (SGLang/vLLM), while matching baselines within noise on workloads with no exploitable prefix structure. Wins hold as KV-cache pressure and inference parallelism scale.

📄 PDF Abstract BibTeX arXiv:2607.02525

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents

2026-05-19 · Zhuohan Gu, Qizheng Zhang, Omar Khattab, Samuel Madden arxiv

Large language model (LLM) agents increasingly operate over long and recurring external contexts, like document corpora and code repositories. Across invocations, existing approaches preserve either the agent's trajector…

A Multi-Modal AI Framework for Real-Time Queue Prediction, Management and Optimisation in Intelligent Border Control Systems

2026-08-27 · Varvara Mama, Eleni Veroni, Nikolaos Kapsalis, Christos D. Nikolopoulos 외 arxiv

In the present work an efficient border control management procedure is proposed. Compared to operational queue management systems, whose operations are based on mostly static data, the proposed work takes into account d…

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live

2025-11-04 · Hanchen Li, Runyuan He, Qiuyang Mang, Qizheng Zhang 외 arxiv

KV cache management is essential for efficient LLM inference. To maximize utilization, existing inference engines evict finished requests' KV cache if new requests are waiting. This policy breaks for agentic workloads, w…

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

2026-05-12 · Shaoke Fang, Ziang Li, Wenfei Wu, Jiatong Ji 외 arxiv

Prefix caching is a key optimization in Large Language Model (LLM) serving, reusing attention Key-Value (KV) states across requests with shared prompt prefixes to reduce expensive prefill computation. However, its benefi…

DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management

2025-12-08 · Zhongchun Zhou, Chengtao Lai, Yuhang Gu, Wei Zhang arxiv

The rapid adoption of large language models (LLMs) is pushing AI accelerators toward increasingly powerful and specialized designs. Instead of further complicating software development with deeply hierarchical scratchpad…