paper-with-me

홈 › Papers

Divergence Timing and Cumulative Disagreement under KV-Cache Eviction

2026-09-15 · Xinyue Luo, Fei Yu arxiv

KV-cache eviction perturbs the conditional token distributions governing autoregressive generation. We investigate how first-divergence timing and subsequent token mismatch determine cumulative disagreement. We derive an exact decomposition under a specified stepwise maximal coupling: the expected mismatch fraction equals a first-mismatch contribution plus post-divergence exposure multiplied by its mismatch rate. An explicit construction over unrestricted autoregressive kernel pairs realizes the sharp interval of risks compatible with a finite divergence-aligned observation window. Residual-branch conditional Monte Carlo provides unbiased joint estimates of occurrence, occupation, and window/tail contributions, with per-replicate variance dominance for total token loss. Complete trajectories from Meta-Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct show that SnapKV at 50% retention enters divergence later and less often than SnapKV-512 or recent-token retention with the same 50% prompt-cache budget, while post-divergence total variation (TV) remains high. In an exploratory analysis of 288 documents, post-divergence exposure accounts for 85-90% of four aggregate mismatch gaps. On 288 independent documents at 90% retention, prespecified comparisons show higher branch-aligned TV in the late than in the early window in both models.

📄 PDF Abstract BibTeX arXiv:2609.16617

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems

2026-04-19 · Yuji Yamamoto, Satoshi Matsuura arxiv

Rowhammer on GPU DRAM has enabled adversarial bit flips in model weights; shared KV-cache blocks in LLM serving systems present an analogous but previously unexamined target. In vLLM's Prefix Caching, these blocks exist …

Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation

2026-07-30 · Alexander Boesgaard Lorup arxiv

Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that originally reached the prefix. We audit that assumption at a whole reasoni…

Auditing Prompt Caching in Language Model APIs

2025-02-11 · Chenchen Gu, Xiang Lisa Li, Rohith Kuditipudi, Percy Liang 외

Prompt caching in large language models (LLMs) results in data-dependent timing variations: cached prompts are processed faster than non-cached prompts. These timing differences introduce the risk of side-channel timing …

DecoderLanguage ModelingLanguage Modellingmodel

Dimensions of Disagreement: Unpacking Divergence and Misalignment in Cognitive Science and Artificial Intelligence

2023-10-03 · Kerem Oktar, Ilia Sucholutsky, Tania Lombrozo, Thomas L. Griffiths

The increasing prevalence of artificial agents creates a correspondingly increasing need to manage disagreements between humans and artificial agents, as well as between artificial agents themselves. Considering this lar…

The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference

2026-04-16 · Ranjith Chodavarapu, Lei Xu arxiv

KV caching is a ubiquitous optimization in autoregressive transformer inference, long presumed to be numerically equivalent to cache-free computation. This assumption fails under standard FP16 precision: cache-ON and cac…