paper-with-me

홈 › Papers

Can I Buy Your KV Cache?

2026-06-11 · Luoyuan Zhang arxiv

Right now, across the world, AI agents are repeating the same absurd act: to read one document, they each recompute it from scratch. Every agent re-runs prefill, the most compute-intensive step a large model takes, over identical text, only to rebuild a key-value (KV) cache identical to the one the agent before it just built. The same answer, computed a million times. We make a proposal that is almost offensively simple: compute it once. Let a publisher precompute a document's KV cache, and let every other agent buy the right to load it and skip prefill. It works, and it is token-exact: loading a precomputed KV and continuing matches prefilling from scratch (24/24 greedy tokens, and at the logits level), with no accuracy cost. On Qwen3-4B, reuse is 9-50x cheaper in compute than prefill, and the gap widens with length (prefill's attention scales with L^2), so a single reuse already pays it back. Then the part that matters: where the KV lives. Shipping it fails, because KV is nearly incompressible, so per-load egress costs more than the prefill it saves. Hosting it provider-side, exactly as production prompt-caching works, removes egress entirely. The size of the prize is set by our measured compute saving: serving one hot 3774-token document to 80M agents costs ~$1.5M to re-prefill but only ~$0.03M of reuse compute (49.7x less). The 0.1x cache-read tariff APIs charge passes a 10x discount to users while sitting inside this measured envelope, so the 10x is a floor that the measured ~50x compute saving clears, and the gap to the physical ~50x is provider margin: millions of dollars per popular document. We frame the resulting agent-native prefill CDN and leave lossless KV compression and a cross-party payment layer as the open problems.

📄 PDF Abstract BibTeX arXiv:2606.13361

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Edit-Your-Interest: Efficient Video Editing via Feature Most-Similar Propagation

2025-10-15 · Yi Zuo, Zitao Wang, Lingling Li, Xu Liu 외 arxiv

Text-to-image (T2I) diffusion models have recently demonstrated significant progress in video editing. However, existing video editing methods are severely limited by their high computational overhead and memory consumpt…

Benchmarking Machine Learning: How Fast Can Your Algorithms Go?

2021-01-08 · Zeyu Ning, Hugues Nelson Iradukunda, Qingquan Zhang, Ting Zhu

This paper is focused on evaluating the effect of some different techniques in machine learning speed-up, including vector caches, parallel execution, and so on. The following content will include some review of the prev…

BenchmarkingBIG-bench Machine Learning

Make Your LVLM KV Cache More Lightweight

2026-05-01 · Xihao Chen, Yangyang Guo, Roger Zimmermann arxiv

Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency in Large Language Models (LLMs), its direct adoption in LVLMs intro…

Harnessing Your DRAM and SSD for Sustainable and Accessible LLM Inference with Mixed-Precision and Multi-level Caching

2024-10-17 · Jie Peng, Zhang Cao, Huaizhi Qu, Zhengyu Zhang 외

Although Large Language Models (LLMs) have demonstrated remarkable capabilities, their massive parameter counts and associated extensive computing make LLMs' deployment the main part of carbon emission from nowadays AI a…

GPUQuantization

Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation

2025-09-20 · Yue Ma, Zexuan Yan, Hongyu Liu, Hongfa Wang 외 arxiv

We present Follow-Your-Emoji-Faster, an efficient diffusion-based framework for freestyle portrait animation driven by facial landmarks. The main challenges in this task are preserving the identity of the reference portr…