paper-with-me

홈 › Papers

Minimal-Intervention KV Retention via Set-Conditioned Diversity

2026-05-14 · Libo Sun, Po-wei Harn, Peixiong He, Xiao Qin arxiv

KV-cache compression at small budgets is a crowded design space spanning cache representation, head-wise routing, compression cadence, decoding behavior, and within-budget scoring. We study seven mechanisms across these five families under matched mean cache on long-form mathematical reasoning (MATH-500~\cite{hendrycks2021math}) with two distilled-reasoning models (Qwen-7B and Llama-8B variants of DeepSeek-R1-Distill~\cite{deepseek2025r1}) at budgets $b \in \{64, 128\}$. All seven were rejected. We then propose $α$, a one-function modification to the TriAttention~\cite{mao2026triattention} retention scorer that replaces argmax-top-$k$ with greedy facility-location-inspired selection under a V-space redundancy penalty controlled by a single weight $λ$. A pre-registered protocol tunes $λ$ on a frozen development split and confirms on a disjoint held-out split; with $λ= 0.5$, $α$ clears Bonferroni on two of the four (model, budget) cells (Qwen $b{=}128$ and Llama $b{=}64$), no cell is significantly negative, and the pre-registered Branch~A triggers. The finding is asymmetric: a minimal scoring modification beat heavier structural redesigns in this regime, and the combined matched-memory, sympy-graded, held-out confirmation protocol is the evidence standard that made the asymmetry visible.

📄 PDF Abstract BibTeX arXiv:2605.14292

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Agentic Control in Variational Language Models

2026-04-14 · Yves Ruffenach arxiv

We study whether a variational language model can support a minimal and measurable form of agentic control grounded in its own internal evidence. Our model combines local variational hidden computation (EVE), a homeostat…

e-Profits: A Business-Aligned Evaluation Metric for Profit-Sensitive Customer Churn Prediction

2025-07-09 · Awais Manzoor, M. Atif Qureshi, Etain Kidney, Luca Longo arxiv

Retention campaigns in customer relationship management often rely on churn prediction models evaluated using traditional metrics such as AUC and F1-score. However, these metrics fail to reflect financial outcomes and ma…

Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks

2025-11-22 · Jie Li, Hongyi Cai, Mingkang Dong, Muxin Pu 외 arxiv

Automatically detecting abnormal events in videos is crucial for modern autonomous systems, yet existing Video Anomaly Detection (VAD) benchmarks lack the scene diversity, balanced anomaly coverage, and temporal complexi…

Video Anomaly DetectionVideo Generation

How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position

2026-06-13 · Yang Yu, Shiyuan Zhang, Yifei Sheng, Haoxiang Ren 외 arxiv

World models have become a central abstraction in modern AI. The term now refers to several different objects: action-conditioned environment models, latent imagination models, future-video predictors, interactive neural…

Instruction FollowingValue prediction

G2: Guided Generation for Enhanced Output Diversity in LLMs

2025-11-01 · Zhiwen Ruan, Yixia Li, Yefeng Liu, Yun Chen 외 arxiv

Large Language Models (LLMs) have demonstrated exceptional performance across diverse natural language processing tasks. However, these models exhibit a critical limitation in output diversity, often generating highly si…