paper-with-me

홈 › Papers

On the Efficacy of Eviction Policy for Key-Value Constrained Generative Language Model Inference

2024-02-09 · Siyu Ren, Kenny Q. Zhu

Despite the recent success associated with Large Language Models (LLMs), they are notably cost-prohibitive to deploy in resource-constrained environments due to their excessive memory and computational demands. In addition to model parameters, the key-value cache is also stored in GPU memory, growing linearly with batch size and sequence length. As a remedy, recent works have proposed various eviction policies for maintaining the overhead of key-value cache under a given budget. This paper embarks on the efficacy of existing eviction policies in terms of importance score calculation and eviction scope construction. We identify the deficiency of prior policies in these two aspects and introduce RoCo, a robust cache omission policy based on temporal attention scores and robustness measures. Extensive experimentation spanning prefilling and auto-regressive decoding stages validates the superiority of RoCo. Finally, we release EasyKV, a versatile software package dedicated to user-friendly key-value constrained generative inference. Code available at https://github.com/DRSY/EasyKV.

📄 PDF Abstract BibTeX arXiv:2402.06262

Code (1)

drsy/easykv 공식 구현 pytorch

Tasks

GPULanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

2026-07-06 · Lukas Hauzenberger, Niklas Schmidinger, Anamaria-Roberta Hartl, David Stap 외 arxiv

Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length. Existing KV eviction methods often rely on static heuristics or proxy scores, whic…

Mathematical Reasoning

Not All Tokens Are Worth Caching: Learning Semantic-Aware Eviction for LLM Prefix Caches

2026-05-12 · Shaoke Fang, Ziang Li, Wenfei Wu, Jiatong Ji 외 arxiv

Prefix caching is a key optimization in Large Language Model (LLM) serving, reusing attention Key-Value (KV) states across requests with shared prompt prefixes to reduce expensive prefill computation. However, its benefi…

Service-Induced Congestion in Memory-Constrained LLM Serving

2026-06-14 · Ruicheng Ao, Jing Dong, Gan Luo, David Simchi-Levi arxiv

In large language model (LLM) serving, each request accumulates persistent graphics processing unit (GPU) memory during service as its key-value cache grows with every generated token. Under high concurrency, aggregate m…

LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning

2025-06-19 · Haoyue Zhang, Hualei Zhang, Xiaosong Ma, Jie Zhang 외

Large Language Models (LLMs) exhibit enhanced reasoning capabilities by employing Chain-of-Thought (CoT). However, the extended reasoning sequences introduce significant GPU memory overhead due to increased key-value (KV…

GPU

D2O: Dynamic Discriminative Operations for Efficient Generative Inference of Large Language Models

2024-06-18 · Zhongwei Wan, Xinjian Wu, Yu Zhang, Yi Xin 외

Efficient inference in Large Language Models (LLMs) is impeded by the growing memory demands of key-value (KV) caching, especially for longer sequences. Traditional KV cache eviction strategies, which prioritize less cri…

Text Generation