paper-with-me

홈 › Papers

DistrAttention: An Efficient and Flexible Self-Attention Mechanism on Modern GPUs

2025-07-23 · Haolin Jin, Mengbai Xiao, Yuan Yuan, Xiao Zhang, Dongxiao Yu, Guanghui Zhang, Haoliang Wang arxiv

The Transformer architecture has revolutionized deep learning, delivering the state-of-the-art performance in areas such as natural language processing, computer vision, and time series prediction. However, its core component, self-attention, has the quadratic time complexity relative to input sequence length, which hinders the scalability of Transformers. The exsiting approaches on optimizing self-attention either discard full-contextual information or lack of flexibility. In this work, we design DistrAttention, an effcient and flexible self-attention mechanism with the full context. DistrAttention achieves this by grouping data on the embedding dimensionality, usually referred to as $d$. We realize DistrAttention with a lightweight sampling and fusion method that exploits locality-sensitive hashing to group similar data. A block-wise grouping framework is further designed to limit the errors introduced by locality sensitive hashing. By optimizing the selection of block sizes, DistrAttention could be easily integrated with FlashAttention-2, gaining high-performance on modern GPUs. We evaluate DistrAttention with extensive experiments. The results show that our method is 37% faster than FlashAttention-2 on calculating self-attention. In ViT inference, DistrAttention is the fastest and the most accurate among approximate self-attention mechanisms. In Llama3-1B, DistrAttention still achieves the lowest inference time with only 1% accuray loss.

📄 PDF Abstract BibTeX arXiv:2507.17245

Code (0)

등록된 구현이 없습니다.

Tasks

Time Series Prediction

Similar Papers 제목 키워드 기반

On the Role of Hidden States of Modern Hopfield Network in Transformer

2025-11-24 · Tsubasa Masumura, Masato Taki arxiv

Associative memory models based on Hopfield networks and self-attention based on key-value mechanisms have been popular approaches in the study of memory mechanisms in deep learning. It has been pointed out that the stat…

PaTH Attention: Position Encoding via Accumulating Householder Transformations

2025-05-22 · Songlin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan 외

The attention mechanism is a core primitive in modern large language models (LLMs) and AI more broadly. Since attention by itself is permutation-invariant, position encoding is essential for modeling structured domains s…

Language ModelingLanguage ModellingPosition

How Powerful Potential of Attention on Image Restoration?

2024-03-15 · Cong Wang, Jinshan Pan, Yeying Jin, Liyan Wang 외

Transformers have demonstrated their effectiveness in image restoration tasks. Existing Transformer architectures typically comprise two essential components: multi-head self-attention and feed-forward network (FFN). The…

Image Restoration

On The Computational Complexity of Self-Attention

2022-09-11 · Feyza Duman Keles, Pruthuvi Mahesakya Wijewardena, Chinmay Hegde

Transformer architectures have led to remarkable progress in many state-of-art applications. However, despite their successes, modern transformers rely on the self-attention mechanism, whose time- and space-complexity is…

Learning to Scale Temperature in Masked Self-Attention for Image Inpainting

2023-02-13 · Xiang Zhou, Yuan Zeng, Yi Gong

Recent advances in deep generative adversarial networks (GAN) and self-attention mechanism have led to significant improvements in the challenging task of inpainting large missing regions in an image. These methods integ…

Image InpaintingPatch Matching