paper-with-me

홈 › Papers

KVSculpt: KV Cache Compression as Distillation

2026-03-29 · Bo Jiang, Sian Jin arxiv

KV cache compression is critical for efficient long-context LLM inference. Approaches that reduce the per-pair footprint -- quantization and low-rank decomposition -- are orthogonal to those that reduce the sequence length of the cache. Along the sequence-length dimension, existing methods range from pure eviction -- selecting which KV pairs to keep -- to merging, which combines similar pairs into fewer ones. Both remain anchored to the original cache entries. We propose KVSculpt, which moves to the other end of this spectrum: instead of selecting or combining original pairs, we optimize a smaller set of unconstrained KV pairs in continuous embedding space to preserve each layer's attention behavior. Keys are optimized via L-BFGS and values are solved in closed form via least squares, alternating every few steps. On top of this, we introduce adaptive budget allocation, which uses a cheap pilot compression run to redistribute the compression budget across layers and KV heads based on per-component difficulty. On Qwen2.5-1.5B-Instruct with 2048-token contexts, KVSculpt reduces KL divergence by 3.5-4.1x compared to Select+Fit -- attention-score eviction with least-squares value fitting -- across compression ratios r in {0.3, 0.5, 0.7}. Adaptive allocation provides an additional 1.3x KL reduction at no extra inference cost. Analysis reveals that compression difficulty is highly non-uniform: per-layer pilot MSE varies by up to 100x across layers, and the two KV heads within a single layer can differ by up to 467x -- demonstrating that fine-grained budget allocation is essential.

📄 PDF Abstract BibTeX arXiv:2603.27819

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons

2025-10-15 · Giovanni Monea, Yair Feldman, Shankar Padmanabhan, Kianté Brantley 외 arxiv

The scalability of large language models for long-context reasoning is severely constrained by the linear growth of their Transformer key-value cache, which incurs significant memory and computational costs. We posit tha…

Reinforcement Learning

MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection

2024-10-16 · Bokai Lin, Zihao Zeng, Zipeng Xiao, Siqi Kou 외

KV cache has become a de facto technique for the inference of large language models (LLMs), where tensors of shape (layer number, head number, sequence length, feature dimension) are introduced to cache historical inform…

Data Compression

G-KV: Decoding-Time KV Cache Eviction with Global Attention

2025-11-29 · Mengqi Liao, Lu Wang, Chaoyun Zhang, Zekai Shen 외 arxiv

Recent reasoning large language models (LLMs) excel in complex tasks but encounter significant computational and memory challenges due to long sequence lengths. KV cache compression has emerged as an effective approach t…

Reinforcement Learning

ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs

2026-05-30 · Yiling Gao, Hongchen Wei, Zhenzhong Chen arxiv

In Vision-Language Models (VLMs), high-resolution images produce a large number of visual tokens, resulting in high computational costs and KV-cache overhead during inference. To address this problem, we propose an Extre…

PANDO: Efficient Multimodal AI Agents via Online Skill Distillation

2026-05-24 · Yubo Li, Yidi Miao, Yuntian Shen, Yuxin Liu arxiv

Recent advances in multimodal web agents often rely on increased inference-time computation, including rollout search, verifier passes, offline skill discovery, and specialist model stacks. This raises a central question…