paper-with-me

홈 › Papers

Position-Agnostic Pre-Projection for Transformer Attention: Nonlinear Feature Construction and Content Skip Before Q/K/V

2026-04-12 · Chirag Shinde arxiv

We propose two complementary modifications to transformer attention blocks. First, a non-linear pre-projection MLP is inserted between layer norm and Q/K/V projections, constructing richer features in a position-agnostic manner before any positional encoding is applied. Second, a content skip connection routes the pre-projection's features around the attention mechanism, allowing content information to bypass position-aware attention where beneficial. In frozen-probe experiments on Pythia-160M and 410M, the combined approach achieves the strongest results across methods: +40.6% LAMBADA accuracy and -39% perplexity at 160M scale. Learned skip connection weights reveal a consistent pattern across model sizes: later transformer layers activate the content bypass more strongly than earlier layers, suggesting that deeper layers benefit from content information that does not pass through positional attention. All modifications add no K/V cache overhead.

📄 PDF Abstract BibTeX arXiv:2604.10791

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FourierQK: Spectral Preprocessing of Query-Key Projections Improves Transformer Attention

2026-07-08 · Athanasios Zeris arxiv

FFT-based spectral preprocessing of learned query-key (Q/K) projections substantially improves transformer attention on character-level language modelling. On TinyShakespeare: a fixed random spectral filter achieves val=…

Language Modelling

Beyond Linearity in Attention Projections: The Case for Nonlinear Queries

2026-03-11 · Marko Karbevski arxiv

Recent algebraic analysis shows that in decoder-only and encoder-only transformers, the Query projection $W_Q$ may be set to identity without noticeable performance deterioration. This is possible because attention depen…

Self-Attention as Distributional Projection: A Unified Interpretation of Transformer Architecture

2025-11-16 · Nihal Mehta arxiv

This paper presents a mathematical interpretation of self-attention by connecting it to distributional semantics principles. We show that self-attention emerges from projecting corpus-level co-occurrence statistics into …

Do Transformers Need Three Projections? Systematic Study of QKV Variants

2026-06-01 · Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis arxiv

Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and …

Accelerating Attention with Basis Decomposition

2025-10-02 · Jialin Zhao arxiv

Attention is a core operation in large language models (LLMs). We present BD Attention (BDA), a lossless algorithmic reformulation of attention. BDA is enabled by a simple matrix identity from Basis Decomposition (BD), w…