paper-with-me

홈 › Papers

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling

2025-05-30 · Xiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin Cui

Many advanced Large Language Model (LLM) applications require long-context processing, but the self-attention module becomes a bottleneck during the prefilling stage of inference due to its quadratic time complexity with respect to sequence length. Existing sparse attention methods accelerate attention computation by skipping less significant regions of the attention map. However, these approaches typically perform coarse-grained inspection of the attention map, rendering considerable loss in model accuracy. In this paper, we propose SALE, a fine-grained sparse attention method that accelerates the long-context prefilling stage of LLM with negligible loss in model accuracy. SALE achieves fast and accurate fine-grained attention weight estimation through 4-bit quantized query-key products, followed by block-sparse attention to accelerate prefilling computations. For importance evaluation for query-key pairs, we adopt our Relative Attention Score metric, which offers significantly higher efficiency within our framework. We implement a custom CUDA kernel optimized for our approach for hardware efficiency, reducing the additional overhead to approximately 11% of the full attention latency. Notably, SALE requires no parameter training and can be seamlessly integrated into existing systems with trivial code modifications. Experiments on long-context benchmarks demonstrate that our method outperforms existing approaches in accuracy-efficiency trade-offs, achieving at least 3.36x speedups on Llama-3.1-8B for sequences longer than 64K while maintaining model quality.

📄 PDF Abstract BibTeX arXiv:2505.24179

Code (1)

birdchristopher/sale 공식 구현 pytorch

Tasks

Large Language Model

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
ADOPT Please enter a description about the method here

Similar Papers 제목 키워드 기반

Double-P: Hierarchical Top-P Sparse Attention for Long-Context LLMs

2026-02-05 · Wentao Ni, Kangqi Zhang, Zhongming Yu, Oren Nelson 외 arxiv

As long-context inference becomes central to large language models (LLMs), attention over growing key-value caches emerges as a dominant decoding bottleneck, motivating sparse attention for scalable inference. Fixed-budg…

MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation

2025-10-21 · Weinan Jia, Yuning Lu, Mengqi Huang, Hualiang Wang 외 arxiv

Long video generation with Diffusion Transformers (DiTs) is bottlenecked by the quadratic scaling of full attention with sequence length. Since attention is highly redundant, outputs are dominated by a small subset of qu…

Video Generation

Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing

2025-05-26 · Dan Peng, Zhihui Fu, Zewen Ye, Zhuoran Song 외

Sparse attention methods exploit the inherent sparsity in attention to speed up the prefilling phase of long-context inference, mitigating the quadratic complexity of full attention computation. While existing sparse att…

Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference

2025-10-21 · Siyuan Yan, Guo-Qing Jiang, Yuchen Zhang, Xiaoxing Ma 외 arxiv

Large language models (LLMs) now support context windows of hundreds of thousands to millions of tokens, enabling applications such as long-document summarization, large-scale code synthesis, multi-document question answ…

Document SummarizationQuestion Answering

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding

2026-06-23 · WenHung Lee, Jian-Jia Chen, Xiaolin Lin, Pei-Shuo Wang 외 arxiv

While speculative decoding improves inference throughput for multi-batch long-context Large Language Models (LLMs), its efficiency is often limited by a verification bottleneck where Key-Value (KV) cache loading dominate…