paper-with-me

홈 › Papers

Block Sparse Attention with Log-Linear Complexity

2026-09-25 · Bohao Tang, Zhen Qin, Yuqi Pan, Zheng Li, Pengfei Liu hf

Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadratic in sequence length. To address this issue, we propose PISA, a block-sparse attention mechanism that employs a pyramid Top-K selection strategy. The main idea is to gradually narrow down the candidates across different levels, making it more efficient to find the most relevant keys. Specifically, we construct a coarse-to-fine hierarchy of keys and perform selection from the coarsest level. At each level, LogSumExp scoring is applied to a bounded candidate set to select candidates for the next finer level, continuing until the finest level is reached. Through pooling, we construct O(log N) levels of keys, yielding an overall complexity of O(Nlog N), where N denotes the sequence length. We develop hardware-aware Triton kernels for both training and inference, fusing hierarchical routing and LogSumExp scoring without materializing the query-key score matrix. We further evaluate our method on language modeling tasks. Compared with the baseline, our method achieves comparable performance on benchmarks such as commonsense reasoning while delivering better results on retrieval tasks.

📄 PDF Abstract BibTeX arXiv:2609.31093

Code (3)

AtharvaDomale/Daily-HuggingFace-AI-Papers ★ 16
qqtang-code/PISA-Project-Page
qqtang-code/Paper-Reading-Collection ★ 1

Similar Papers 제목 키워드 기반

VSANet: View-aware Sparse Attention Network for Light Field Image Denoising

2026-06-23 · Gargi Panda, Soumitra Kundu, Saumik Bhattacharya, Aurobinda Routray arxiv

Light field (LF) image denoising is challenging due to the high-dimensional structure of LF data. While noise is independent across sub-aperture images, scene content exhibits strong cross-view correlations. We introduce…

Image Denoising

SPLA: Block Sparse Plus Linear Attention for Long Context Modeling

2026-01-29 · Bailin Wang, Dan Friedman, Tao Lei, Chong Wang arxiv

Block-wise sparse attention offers significant efficiency gains for long-context modeling, yet existing methods often suffer from low selection fidelity and cumulative contextual loss by completely discarding unselected …

Continual PretrainingGeneral Knowledge

PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers

2026-02-01 · Haopeng Li, Shitong Shao, Wenliang Zhong, Zikai Zhou 외 arxiv

Diffusion Transformers are fundamental for video and image generation, but their efficiency is bottlenecked by the quadratic complexity of attention. While block sparse attention accelerates computation by attending only…

Image Generation

Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers

2025-12-18 · Yifan Zhou, Zeqi Xiao, Tianyi Wei, Shuai Yang 외 arxiv

Diffusion Transformers (DiTs) set the state of the art in visual generation, yet their quadratic self-attention cost fundamentally limits scaling to long token sequences. Recent Top-K sparse attention approaches reduce t…

Image Generation

Sparser Block-Sparse Attention via Token Permutation

2025-10-24 · Xinghao Wang, Pengyu Wang, Dong Zhang, Chenkun Tan 외 arxiv

Scaling the context length of large language models (LLMs) offers significant benefits but is computationally expensive. This expense stems primarily from the self-attention mechanism, whose $O(N^2)$ complexity with resp…

Computational Efficiency