paper-with-me

홈 › Papers

A Method for Building Large Language Models with Predefined KV Cache Capacity

2024-11-24 · Zhonghua Yi, Ge Niu, Lei Wang, Wei Tang, Liqiu Zhang

This paper introduces a novel approach, the Bounded-Cache Transformer (BCT), for building large language models with a predefined Key-Value (KV) cache capacity. The BCT addresses the excessive memory consumption issue in traditional KV caches by implementing a bounded-length KV cache, which is particularly suitable for the attention layers in Transformer decode-only architectures. By dynamically updating the key-value vector sequences, the BCT achieves efficient inference within limited cache capacity, significantly reducing memory usage while maintaining model performance and system throughput. Experimental results demonstrate that the BCT significantly reduces memory usage while maintaining the model's inference quality, offering a new solution for efficient inference in large language models.

📄 PDF Abstract BibTeX arXiv:2411.15785

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Adam 설명 없음
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

QCQA: Quality and Capacity-aware grouped Query Attention

2024-06-08 · Vinay Joshi, Prashant Laddha, Shambhavi Sinha, Om Ji Omer 외

Excessive memory requirements of key and value features (KV-cache) present significant challenges in the autoregressive inference of large language models (LLMs), restricting both the speed and length of text generation.…

Text Generation

Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System

2025-08-17 · Yunhua Fang, Rui Xie, Asad Ul Haq, Linsen Ma 외 arxiv

Large Language Model (LLM) inference is increasingly constrained by memory bandwidth, with frequent access to the key-value (KV) cache dominating data movement. While attention sparsity reduces some memory traffic, the r…

RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression

2025-02-19 · Payman Behnam, Yaosheng Fu, Ritchie Zhao, Po-An Tsai 외

Transformer-based Large Language Models rely critically on KV cache to efficiently handle extended contexts during the decode phase. Yet, the size of the KV cache grows proportionally with the input length, burdening bot…

GPU

AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents

2026-04-27 · Hojoon Kim, Yuheng Wu, Thierry Tambe arxiv

Embodied AI agents increasingly rely on large language models (LLMs) for planning, yet per-step LLM calls impose severe latency and cost. In this paper, we show that embodied tasks exhibit strong plan locality, where the…

CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion

2024-05-26 · Jiayi Yao, Hanchen Li, YuHan Liu, Siddhant Ray 외

Large language models (LLMs) often incorporate multiple text chunks in their inputs to provide the necessary contexts. To speed up the prefill of the long LLM inputs, one can pre-compute the KV cache of a text and re-use…

Language ModelingLanguage ModellingLarge Language ModelRAG