paper-with-me

홈 › Papers

Attention as an RNN

2024-05-22 · Leo Feng, Frederick Tung, Hossein Hajimirsadeghi, Mohamed Osama Ahmed, Yoshua Bengio, Greg Mori

The advent of Transformers marked a significant breakthrough in sequence modelling, providing a highly performant architecture capable of leveraging GPU parallelism. However, Transformers are computationally expensive at inference time, limiting their applications, particularly in low-resource settings (e.g., mobile and embedded devices). Addressing this, we (1) begin by showing that attention can be viewed as a special Recurrent Neural Network (RNN) with the ability to compute its \textit{many-to-one} RNN output efficiently. We then (2) show that popular attention-based models such as Transformers can be viewed as RNN variants. However, unlike traditional RNNs (e.g., LSTMs), these models cannot be updated efficiently with new tokens, an important property in sequence modelling. Tackling this, we (3) introduce a new efficient method of computing attention's \textit{many-to-many} RNN output based on the parallel prefix scan algorithm. Building on the new attention formulation, we (4) introduce \textbf{Aaren}, an attention-based module that can not only (i) be trained in parallel (like Transformers) but also (ii) be updated efficiently with new tokens, requiring only constant memory for inferences (like traditional RNNs). Empirically, we show Aarens achieve comparable performance to Transformers on $38$ datasets spread across four popular sequential problem settings: reinforcement learning, event forecasting, time series classification, and time series forecasting tasks while being more time and memory-efficient.

📄 PDF Abstract BibTeX arXiv:2405.13956

Code (1)

claCase/Attention-as-RNN tf

Tasks

GPUTime SeriesTime Series ClassificationTime Series Forecasting

Similar Papers 제목 키워드 기반

SageAttention2++: A More Efficient Implementation of SageAttention2

2025-05-27 · Jintao Zhang, Xiaoming Xu, Jia Wei, Haofeng Huang 외

The efficiency of attention is critical because its time complexity grows quadratically with sequence length. SageAttention2 addresses this by utilizing quantization to accelerate matrix multiplications (Matmul) in atten…

QuantizationVideo Generation

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse

2026-02-01 · Zizhuo Fu, Wenxuan Zeng, Runsheng Wang, Meng Li arxiv

Large Language Models (LLMs) often assign disproportionate attention to the first token, a phenomenon known as the attention sink. Several recent approaches aim to address this issue, including Sink Attention in GPT-OSS …

Power-based Partial Attention: Bridging Linear-Complexity and Full Attention

2026-01-24 · Yufeng Huang arxiv

It is widely accepted from transformer research that "attention is all we need", but the amount of attention required has never been systematically quantified. Is quadratic $O(L^2)$ attention necessary, or is there a sub…

Flex Attention: A Programming Model for Generating Optimized Attention Kernels

2024-12-07 · Juechu Dong, Boyuan Feng, Driss Guessous, Yanbo Liang 외

Over the past 7 years, attention has become one of the most important primitives in deep learning. The primary approach to optimize attention is FlashAttention, which fuses the operation together, drastically improving b…

Person Re-identification via Attention Pyramid

2021-08-11 · Guangyi Chen, Tianpei Gu, Jiwen Lu, Jin-An Bao 외

In this paper, we propose an attention pyramid method for person re-identification. Unlike conventional attention-based methods which only learn a global attention map, our attention pyramid exploits the attention region…

Person Re-Identification