paper-with-me

홈 › Papers

A Reproduction Study: The Kernel PCA Interpretation of Self-Attention Fails Under Scrutiny

2025-05-12 · Karahan Sarıtaş, Çağatay Yıldız

In this reproduction study, we revisit recent claims that self-attention implements kernel principal component analysis (KPCA) (Teo et al., 2024), positing that (i) value vectors $V$ capture the eigenvectors of the Gram matrix of the keys, and (ii) that self-attention projects queries onto the principal component axes of the key matrix $K$ in a feature space. Our analysis reveals three critical inconsistencies: (1) No alignment exists between learned self-attention value vectors and what is proposed in the KPCA perspective, with average similarity metrics (optimal cosine similarity $\leq 0.32$, linear CKA (Centered Kernel Alignment) $\leq 0.11$, kernel CKA $\leq 0.32$) indicating negligible correspondence; (2) Reported decreases in reconstruction loss $J_\text{proj}$, arguably justifying the claim that the self-attention minimizes the projection error of KPCA, are misinterpreted, as the quantities involved differ by orders of magnitude ($\sim\!10^3$); (3) Gram matrix eigenvalue statistics, introduced to justify that $V$ captures the eigenvector of the gram matrix, are irreproducible without undocumented implementation-specific adjustments. Across 10 transformer architectures, we conclude that the KPCA interpretation of self-attention lacks empirical support.

📄 PDF Abstract BibTeX arXiv:2505.07908

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Unified Geometric Field Theory Framework for Transformers: From Manifold Embeddings to Kernel Modulation

2025-11-11 · Xianshuai Shi, Jianfeng Zhu, Leibo Liu arxiv

The Transformer architecture has achieved tremendous success in natural language processing, computer vision, and scientific computing through its self-attention mechanism. However, its core components-positional encodin…

Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics

2026-05-28 · Matthew Smart, Soumya Ganguly, Nilava Metya, Alexandre V. Morozov 외 arxiv

We study minimal attention-only transformers under all-token corruption and show they admit a two-stage empirical Bayes interpretation. A single attention step computes a kernel-weighted posterior mean with respect to th…

KERPLE: Kernelized Relative Positional Embedding for Length Extrapolation

2022-05-20 · Ta-Chung Chi, Ting-Han Fan, Peter J. Ramadge, Alexander I. Rudnicky

Relative positional embeddings (RPE) have received considerable attention since RPEs effectively model the relative distance among tokens and enable length extrapolation. We propose KERPLE, a framework that generalizes r…

DiversityLanguage ModelingLanguage ModellingPosition

An Inductive Formalization of Self Reproduction in Dynamical Hierarchies

2018-06-23 · Janardan Misra

Formalizing self reproduction in dynamical hierarchies is one of the important problems in Artificial Life (AL) studies. We study, in this paper, an inductively defined algebraic framework for self reproduction on macros…

Artificial Life

Matter-antimatter asymmetry restrains the dimensionality of neural representations: quantum decryption of large-scale neural coding

2022-06-17 · Sofia Karamintziou

Projections from the study of the human universe onto the study of the self-organizing brain are herein leveraged to address certain concerns raised in latest neuroscience research, namely (i) the extent to which neural …

Disentanglement