paper-with-me

홈 › Papers

A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference

2024-10-18 · You Wu, HaoYi Wu, Kewei Tu

Recently, sharing key-value (KV) cache across layers has been found effective in efficient inference of large language models (LLMs). To systematically investigate different techniques of cross-layer KV sharing, we propose a unified framework that covers several recent methods and their novel variants. We conduct comprehensive experiments on all the configurations of the framework, evaluating their generation throughput and performance in language modeling and downstream tasks. We find that when reducing the size of the KV cache by 2x, most configurations can achieve competitive performance to and higher throughput than standard transformers, but when further reducing the size of the KV cache, pairing queries of all layers with KVs of upper layers can better maintain performance, although it also introduces additional training cost and prefilling latency. We hope that this work will help users choose the appropriate approach according to their requirements and facilitate research on the acceleration of LLM inference.

📄 PDF Abstract BibTeX arXiv:2410.14442

Code (1)

whyNLP/LCKV 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Rethinking Parameter Sharing as Graph Coloring for Structured Compression

2025-11-10 · Boyang Zhang, Daning Cheng, Yunquan Zhang arxiv

Modern deep models have massive parameter sizes, leading to high inference-time memory usage that limits practical deployment. Parameter sharing, a form of structured compression, effectively reduces redundancy, but exis…

EPAS: Efficient Training with Progressive Activation Sharing

2026-01-27 · Rezaul Karim, Maryam Dialameh, Yang Liu, Boxing Chen 외 arxiv

We present a novel method for Efficient training with Progressive Activation Sharing (EPAS). This method bridges progressive training paradigm with the phenomenon of redundant QK (or KV ) activations across deeper layers…

Continual Pretraining

Uncovering the Limits of Proof Sharing for Neural Networks

2026-08-19 · Kanak Das, Shubham Ugare, Bor-Yuh Evan Chang, Sasa Misailovic 외 arxiv

Robustness verification of neural networks is increasingly important, due to their use in many critical domains. In certain scenarios, proof sharing has been shown to accelerate incomplete verification techniques by reus…

Basis Sharing: Cross-Layer Parameter Sharing for Large Language Model Compression

2024-10-02 · Jingcun Wang, Yu-Guang Chen, Ing-Chao Lin, Bing Li 외

Large Language Models (LLMs) have achieved remarkable breakthroughs. However, the huge number of parameters in LLMs require significant amount of memory storage in inference, which prevents their practical deployment in …

Language ModelingLanguage ModellingLarge Language ModelModel Compression

DynaShare: Task and Instance Conditioned Parameter Sharing for Multi-Task Learning

2023-05-26 · Elahe Rahimian, Golara Javadi, Frederick Tung, Gabriel Oliveira

Multi-task networks rely on effective parameter sharing to achieve robust generalization across tasks. In this paper, we present a novel parameter sharing method for multi-task learning that conditions parameter sharing …

Multi-Task Learning