paper-with-me

홈 › Papers

KV Admission: Learning What to Write for Efficient Long-Context Inference

2025-12-19 · Yen-Chieh Huang, Pi-Cheng Hsiu, Rui Fang, Ming-Syan Chen arxiv

Long-context LLM inference is bottlenecked by the quadratic attention complexity and linear KV cache growth. Prior approaches mitigate this via post-hoc selection or eviction but overlook the root inefficiency: indiscriminate writing to memory. In this paper, we formalize KV cache management as a causal system of three primitives: KV Admission, Selection, and Eviction. We instantiate KV Admission via Write-Gated KV (WG-KV), a lightweight mechanism that learns to predict token utility before cache entry. By filtering out low-utility states early to maintain a compact global cache alongside a sliding local cache, WG-KV reduces memory usage by 46-68% and delivers 3.03-3.70x prefill and 1.85-2.56x decode speedups on Llama and Qwen models, while maintaining compatibility with FlashAttention and Paged-KV systems. These results demonstrate that learning what to write is a principled and practical recipe for efficient long-context inference. Code is available at https://github.com/EMCLab-Sinica/WG-KV.

📄 PDF Abstract BibTeX arXiv:2512.17452

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

2026-07-25 · Yan Zhang, Shibo Li arxiv

LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated fact written at one step therefore persists as a false premise for …

MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents

2026-05-01 · Tianyu Hu, Weikai Lin, Weizhi Zhang, Jing Ma 외 arxiv

Long-term conversational agents must decide which turns to store in external memory, yet recent systems rely on autoregressive LLM generation at every turn to make that decision. We present MemRouter, a write-side memory…

Answer Generation

Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays

2025-03-25 · Jinsook Lee, AJ Alvero, Thorsten Joachims, René Kizilcec

People are increasingly using technologies equipped with large language models (LLM) to write texts for formal communication, which raises two important questions at the intersection of technology and society: Who do LLM…

Erase-then-Delta Attention: Decoupling Erase and Write Addresses in Delta-Rule Linear Attention

2026-06-25 · Xiao Li, Chengruidong Zhang, Hao Luo, Xi Lin 외 arxiv

Delta-rule linear attention improves recurrent memory updates by correcting what is already stored at the current write address before writing new content. However, the active correction is still anchored to that same wr…

WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems

2026-05-23 · Jiangnan Yu, Kisson Songqi Lin, Jilong Wu arxiv

Long-context memory systems often fail under fixed budgets, but end-to-end evaluation does not reveal whether evidence was discarded during compression or preserved but never retrieved. We introduce a four-condition diag…