paper-with-me

홈 › Papers

Attention-based PCA

2026-05-18 · Rodrigo Maulen-Soto, Claire Boyer arxiv

We study attention mechanisms through the lens of a canonical unsupervised problem: principal component analysis (PCA). We show that, when trained on Gaussian data, both softmax and linear attention layers learn parameters that align with the principal eigenvectors of the covariance matrix, thereby establishing a direct and explicit connection with PCA. Our analysis covers both finite and infinite prompt regimes. In the infinite-prompt limit, we prove convergence to globally optimal solutions aligned with the leading spectral direction, while in the finiteprompt setting we show that the same behavior emerges up to sampling effects. We further extend the analysis to an in-context setting with spiked Wishart covariances, where attention successfully recovers the underlying signal direction. These results demonstrate that attention inherently performs PCA-like computations under unsupervised objectives, providing a theoretical foundation for its representation-learning capabilities.

📄 PDF Abstract BibTeX arXiv:2605.18315

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SageAttention2++: A More Efficient Implementation of SageAttention2

2025-05-27 · Jintao Zhang, Xiaoming Xu, Jia Wei, Haofeng Huang 외

The efficiency of attention is critical because its time complexity grows quadratically with sequence length. SageAttention2 addresses this by utilizing quantization to accelerate matrix multiplications (Matmul) in atten…

QuantizationVideo Generation

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse

2026-02-01 · Zizhuo Fu, Wenxuan Zeng, Runsheng Wang, Meng Li arxiv

Large Language Models (LLMs) often assign disproportionate attention to the first token, a phenomenon known as the attention sink. Several recent approaches aim to address this issue, including Sink Attention in GPT-OSS …

Power-based Partial Attention: Bridging Linear-Complexity and Full Attention

2026-01-24 · Yufeng Huang arxiv

It is widely accepted from transformer research that "attention is all we need", but the amount of attention required has never been systematically quantified. Is quadratic $O(L^2)$ attention necessary, or is there a sub…

Flex Attention: A Programming Model for Generating Optimized Attention Kernels

2024-12-07 · Juechu Dong, Boyuan Feng, Driss Guessous, Yanbo Liang 외

Over the past 7 years, attention has become one of the most important primitives in deep learning. The primary approach to optimize attention is FlashAttention, which fuses the operation together, drastically improving b…

Person Re-identification via Attention Pyramid

2021-08-11 · Guangyi Chen, Tianpei Gu, Jiwen Lu, Jin-An Bao 외

In this paper, we propose an attention pyramid method for person re-identification. Unlike conventional attention-based methods which only learn a global attention map, our attention pyramid exploits the attention region…

Person Re-Identification