paper-with-me

홈 › Papers

UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training

2026-05-26 · Keqi Deng, Shaoshi Ling, Ruchao Fan, Jinyu Li arxiv

Long-context inference in large language models (LLMs) is bottlenecked by the linear growth of the self-attention key-value (KV) cache. Top-k sparse attention alleviates this by loading only a small fraction of the KV cache, but accurately and cheaply estimating cache importance, for both training-free use and sparsity-aware training, remains challenging. This paper proposes UNIQUE, a universal top-k sparse attention framework that addresses both requirements and stays consistently effective across LLM modalities. UNIQUE operates at the granularity of KV pages and estimates per-page importance with a simple yet accurate score combining the mean of the page's keys as a representative vector with their standard deviation as an offset term. To further close the train-inference gap, this paper introduces a soft-mask sparsity-aware training scheme that uses the top-k score boundary as a per-query threshold and a sigmoid soft mask around it, requiring neither auxiliary losses nor architectural changes. Experiments on text and speech LLMs show that UNIQUE preserves task performance on long-context benchmarks such as LongBench Pro and on long-form speech recognition, while delivering up to 11.4x attention-kernel speedup over FlashInfer dense attention and at least 5.3x end-to-end decoding speedup over a vLLM-based dense model.

📄 PDF Abstract BibTeX arXiv:2605.27740

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference

2025-02-25 · Jintao Zhang, Chendong Xiang, Haofeng Huang, Jia Wei 외

An efficient attention implementation is essential for large models due to its quadratic time complexity. Fortunately, attention commonly exhibits sparsity, i.e., many values in the attention map are near zero, allowing …

modelVideo Generation

When can dictionary learning uniquely recover sparse data from subsamples?

2015-07-31

Sparse coding or sparse dictionary learning has been widely used to recover underlying structure in many kinds of natural data. Here, we provide conditions guaranteeing when this recovery is universal; that is, when spar…

Dictionary LearningLEMMA

Geometry-Aware Attention Guidance for Diffusion Models via Modern Hopfield Dynamics

2026-03-03 · Kwanyoung Kim arxiv

Classifier-Free Guidance (CFG) improves sample quality in diffusion models, but its dual-pass inference and reliance on null-condition training limit its use in few-step regimes. Attention-space guidance has emerged as a…

The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs

2025-04-24 · Piotr Nawrot, Robert Li, Renjie Huang, Sebastian Ruder 외

Sparse attention offers a promising strategy to extend long-context capabilities in Transformer LLMs, yet its viability, its efficiency-accuracy trade-offs, and systematic scaling studies remain unexplored. To address th…

Attention that does not Explain Away

2020-09-29 · Nan Ding, Xinjie Fan, Zhenzhong Lan, Dale Schuurmans 외

Models based on the Transformer architecture have achieved better accuracy than the ones based on competing architectures for a large set of tasks. A unique feature of the Transformer is its universal application of a se…