paper-with-me

홈 › Papers

ENA: Efficient N-dimensional Attention

2025-08-16 · Yibo Zhong arxiv

Efficient modeling of long sequences of high-order data requires a more efficient architecture than Transformer. In this paper, we investigate two key aspects of extending linear recurrent models, especially those originally designed for language modeling, to high-order data (1D to ND): scanning strategies and attention-hybrid architectures. Empirical results suggest that scanning provides limited benefits, while attention-hybrid models yield promising results. Focusing on the latter, we further evaluate types of attention and find that tiled high-order sliding window attention (SWA) is efficient in both theory and practice. We term the resulting hybrid architecture of linear recurrence and high-order SWA as Efficient N-dimensional Attention (ENA). We then conduct several experiments to demonstrate its effectiveness. The intuition behind ENA is that linear recurrence compresses global information into a state, while SWA complements it by enforcing strict local modeling. Together, they form a simple framework that offers a promising and practical solution for ultra-long high-order data modeling.

📄 PDF Abstract BibTeX arXiv:2508.11921

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Your Local GAN: Designing Two Dimensional Local Attention Mechanisms for Generative Models

2019-11-27 · CVPR 2020 6 · Giannis Daras, Augustus Odena, Han Zhang, Alexandros G. Dimakis

We introduce a new local sparse attention layer that preserves two-dimensional geometry and locality. We show that by just replacing the dense attention layer of SAGAN with our construction, we obtain very significant FI…

Conditional Image GenerationDeep AttentionImage Generation

Customizing the Inductive Biases of Softmax Attention using Structured Matrices

2025-09-09 · Yilun Kuang, Noah Amsel, Sanae Lotfi, Shikai Qiu 외 arxiv

The core component of attention is the scoring function, which transforms the inputs into low-dimensional queries and keys and takes the dot product of each pair. While the low-dimensional projection improves efficiency,…

Loki: Low-rank Keys for Efficient Sparse Attention

2024-06-04 · Prajwal Singhania, Siddharth Singh, Shwai He, Soheil Feizi 외

Inference on large language models (LLMs) can be expensive in terms of the compute and memory costs involved, especially when long sequence lengths are used. In particular, the self-attention mechanism used in LLM infere…

PLS in the Mirror of Self-Attention

2026-05-27 · Jiangsheng, You arxiv

This note provides an interesting observation on casting partial least square (PLS) as a linearized self-attention so that PLS may be studied within the neural network paradigm. On the other hand, the dimensionality redu…

Dimensionality Reduction

Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning

2025-08-23 · Junxuan Wang, Xuyang Ge, Wentao Shu, Zhengfu He 외 arxiv

Transformer architectures, and their attention mechanisms in particular, form the foundation of modern large language models. While transformer models are widely believed to operate in high-dimensional hidden spaces, we …