paper-with-me

홈 › Papers

Power-based Partial Attention: Bridging Linear-Complexity and Full Attention

2026-01-24 · Yufeng Huang arxiv

It is widely accepted from transformer research that "attention is all we need", but the amount of attention required has never been systematically quantified. Is quadratic $O(L^2)$ attention necessary, or is there a sub-quadratic attention mechanism that can achieve comparable performance? To answer this question, we introduce power-based partial attention (PPA), an attention mechanism of order $O(L^{1+p})$, where $0 \leq p \leq 1$, such that $p=0$ corresponds to sliding window attention with linear complexity, and $p=1$ corresponds to full attention. With this attention construction, we can explore how transformer architecture performance varies as a function of the attention scaling behavior controlled by $p$. The overall trend from our experiments shows an S-curve-like behavior where the performance transitions from sliding-window (linear-complexity) attention to full attention over a narrow window of $p$ values, and plateaus as $p$ approaches $1$. In our experiments, we show that there exists $0<p<1$ such that $O(L^{1+p})$ attention is sufficient to achieve similar results as $O(L^2)$ full attention.

📄 PDF Abstract BibTeX arXiv:2601.17334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficient High-Accuracy PDEs Solver with the Linear Attention Neural Operator

2025-10-19 · Ming Zhong, Zhenya Yan arxiv

Neural operators offer a powerful data-driven framework for learning mappings between function spaces, in which the transformer-based neural operator architecture faces a fundamental scalability-accuracy trade-off: softm…

Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective

2025-07-02 · Yuxin Mao, Zhen Qin, Jinxing Zhou, Hui Deng 외 arxiv

Autoregressive (AR) models have garnered significant attention in image generation for their ability to effectively capture both local and global structures within visual data. However, prevalent AR models predominantly …

Computational EfficiencyImage Generation

Reinforcement Learning from Partial Observation: Linear Function Approximation with Provable Sample Efficiency

2022-04-20 · Qi Cai, Zhuoran Yang, Zhaoran Wang

We study reinforcement learning for partially observed Markov decision processes (POMDPs) with infinite observation and state spaces, which remains less investigated theoretically. To this end, we make the first attempt …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Bridging the Divide: Reconsidering Softmax and Linear Attention

2024-12-09 · Dongchen Han, Yifan Pu, Zhuofan Xia, Yizeng Han 외

Widely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when dealing with high-resolution inputs. In c…

Towards End-to-End Generative Modeling of Long Videos with Memory-Efficient Bidirectional Transformers

2023-03-20 · CVPR 2023 1 · Jaehoon Yoo, Semin Kim, Doyup Lee, Chiheon Kim 외

Autoregressive transformers have shown remarkable success in video generation. However, the transformers are prohibited from directly learning the long-term dependency in videos due to the quadratic complexity of self-at…

Video Generation