paper-with-me

홈 › Papers

Cartridges at Scale: Training Modular KV Caches over Large Document Collections

2026-06-03 · Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert arxiv

Large Language Models can reason over long contexts, yet prefilling millions of tokens is wasteful as much of the content remains static across queries. Cartridges address this by distilling document collections into reusable key-value (KV) caches that eliminate prefilling while preserving accuracy. A critical limitation of this approach is that cartridges are monolithic and non-compositional: encoding an entire collection into a single KV block does not scale, and naively mixing cartridges trained in isolation collapses performance to near chance. We introduce Cartridges at Scale (CAS), a training framework for scalable multi-cartridge learning with dynamic distractor mixing and a memory-efficient budget manager that rotates hundreds of per-document cartridges between GPU and persistent storage. Our approach scales to collections exceeding a million tokens, improving over a monolithic cartridge by 10-31 points at comparable token budgets. Oracle cartridge accuracy falls within 2-6 points of full in-context learning even at high compression. When paired with retrieval for cartridge selection, CAS matches or exceeds conventional RAG accuracy while consuming 3-4x fewer prompt tokens.

📄 PDF Abstract BibTeX arXiv:2606.04557

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Where Should a Document Live: Context, Representations, or Parameters?

2026-09-15 · Nathanaël Carraz Rakotonirina, Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert arxiv

To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or inje…

Fast KV Compaction via Attention Matching

2026-02-18 · Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim arxiv

Scaling language models to long contexts is often bottlenecked by the size of the key-value (KV) cache. In deployed settings, long contexts are typically managed through compaction in token space via summarization. Howev…

Optimal Web-Scale Tiering as a Flow Problem

2010-12-01 · NeurIPS 2010 12 · Gilbert Leung, Novi Quadrianto, Kostas Tsioutsiouliklis, Alex J. Smola

We present a fast online solver for large scale maximum-flow problems as they occur in portfolio optimization, inventory management, computer vision, and logistics. Our algorithm solves an integer linear program in an on…

ManagementPortfolio Optimization

ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management

2026-01-29 · Zaifeng Pan, Yipeng Shen, Zhengding Hu, Zhuang Wang 외 arxiv

LLM-based multi-agent simulations are increasingly adopted across application domains, but remain difficult to scale due to GPU memory pressure. Each agent maintains private GPU-resident states, including models, prefix …

LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference

2025-10-08 · Yuhan Liu, Yihua Cheng, Jiayi Yao, Yuwei An 외 arxiv

KV cache has traditionally been stored in GPU memory to accelerate the decoding phase of large language model (LLM) inference. However, it is increasingly necessary to move KV caches outside GPU devices, to enable cache …

Question Answering