paper-with-me

Papers

Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA

2025-11-26 · Allison Li, Kristjan Greenewald, Thomas Parnell, Navid Azizan arxiv

Modern large language model (LLM) systems increasingly rely on multi-turn pipelines that are composed of multiple task-specific adapters, yet existing serving frameworks remain inefficient, incurring substantial recomputation overhead when switching between adapters. We present the first LLM serving engine that supports cross-model prefix cache reuse between base and adapted models via Activated LoRA (aLoRA), enabling efficient and fine-grained adapter switching during inference. Our design extends the vLLM framework by introducing base-aligned block hashing and activation-aware masking within the model execution path, permitting cache reuse across models while preserving compatibility with existing serving engine optimizations. Integrated into a production-grade inference stack, this approach supports dynamic adapter activation without excessive key-value tensor recomputation. Evaluation across representative multi-turn, multi-adapter pipelines demonstrates up to 58x end-to-end latency reduction and over 100x time-to-first-token improvement relative to standard LoRA baselines, with benefits that scale with model size and sequence length and manifest across all stages of the request lifecycle. This work bridges parameter-efficient model adaptation with high-performance serving, providing the first complete realization of cross-model KV-cache reuse in modern LLM inference engines.

📄 PDF Abstract BibTeX arXiv:2512.17910

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs

2026-09-15 · Dushyant Rajput arxiv

A common small-model deployment runs one shared backbone with several LoRA specialists that answer over the same context. Serving them naively re-prefills that shared context once per specialist. We study a narrow, pract…

Arithmetic Reasoning

KVShareArena: KV-Cache Reuse Across Contexts and Model Checkpoints

2026-09-09 · Xi Shi, Qian Lou arxiv

LLM serving systems already reuse KV caches, but only when the reused text sits at the very start of the prompt. Two growing workloads break this condition: a retrieval-augmented generation server assembles a different s…

ICaRus: Identical Cache Reuse for Efficient Multi Model Inference

2026-02-27 · Sunghyeon Woo, Jaeeun Kil, Hoseung Kim, Minsub Kim 외 arxiv

Multi model inference has recently emerged as a prominent paradigm, particularly in the development of agentic AI systems. However, in such scenarios, each model must maintain its own Key-Value (KV) cache for the identic…

RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention

2026-06-04 · Yang Liu, ZhaoKai Luo, HuaYi Jin, ZhiYong Wang 외 arxiv

As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastructure. It limits GPU memory capacity, serving concurrency, cache reuse, and distri…

KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider

2025-06-03 · Jiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen 외

Serving large language models (LLMs) is important for cloud providers, and caching intermediate results (KV\$) after processing each request substantially improves serving throughput and latency. However, there is limite…