paper-with-me

홈 › Papers

Still: Amortized KV Cache Compaction in a Single Forward Pass

2026-06-05 · Charles O'Neill, Alex Sandomirsky, Harry Partridge, Mudith Jayasekara, Max Kirkby arxiv

The KV cache is the memory bottleneck of long-horizon language model deployment. Practically, a deployable compactor must be lightweight enough to call during inference, expressive enough to preserve context under constraint, and reusable across a trajectory. Existing compaction methods satisfy only part of this requirement: selection methods are lightweight but subset-bound, while synthesis methods are expressive but rely on per-context optimization. Here we introduce Still, a small per-layer Perceiver trained once against a frozen base model that produces compact keys and values in a single forward pass. On Qwen and Gemma models, Still occupies the favorable side of the speed--quality frontier across compression ratios from $8\times$ to $200\times$ and context lengths from $8$k to $128$k. On the long-context RULER grid, Still exceeds the strongest baseline by 8--22 points. The same compact cache also supports free-form summarization, preserving most of the full-context gain on HELMET and winning a pairwise LongBench summarization comparison against KV-Distill. Because compaction is a forward pass, Still can be applied iteratively, entering a long-horizon regime unavailable to per-context methods. We show that amortization makes long-context cache compaction tractable, and synthesis makes its compact state useful at extreme compression.

📄 PDF Abstract BibTeX arXiv:2606.07878

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study

2026-08-02 · Yujian Liu, Jiabao Ji, Li An, Rohit Jain 외 arxiv

LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck. KV cache compaction can reduce this cost, but most prior methods assume …

Fast KV Compaction via Attention Matching

2026-02-18 · Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim arxiv

Scaling language models to long contexts is often bottlenecked by the size of the key-value (KV) cache. In deployed settings, long contexts are typically managed through compaction in token space via summarization. Howev…

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents

2026-07-09 · Ashwin Gerard Colaco, Nada Lahjouji arxiv

Large language models, and the agents built on them, spend an ever-growing share of their compute and memory on remembering: caching attention keys and values, carrying long prompts, maintaining recurrent state, and stor…

ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models

2025-10-01 · Akshat Ramachandran, Marina Neseem, Charbel Sakr, Rangharajan Venkatesan 외 arxiv

The long-output context generation of large reasoning models enables extended chain of thought (CoT) but also drives rapid growth of the key-value (KV) cache, quickly overwhelming GPU memory. To address this challenge, w…

AI Guardrail Survival under Single-Cycle Agentic Self-Summarization

2026-08-11 · Ted Kwartler, Alan Aqrawi, Arian Abbasi arxiv

Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing safety constraint during compaction drives behavioral violations…