paper-with-me

홈 › Papers

On the Emergence of Position Bias in Transformers

2025-02-04 · Xinyi Wu, Yifei Wang, Stefanie Jegelka, Ali Jadbabaie

Recent studies have revealed various manifestations of position bias in transformer architectures, from the "lost-in-the-middle" phenomenon to attention sinks, yet a comprehensive theoretical understanding of how attention masks and positional encodings shape these biases remains elusive. This paper introduces a novel graph-theoretic framework to analyze position bias in multi-layer attention. Modeling attention masks as directed graphs, we quantify how tokens interact with contextual information based on their sequential positions. We uncover two key insights: First, causal masking inherently biases attention toward earlier positions, as tokens in deeper layers attend to increasingly more contextualized representations of earlier tokens. Second, we characterize the competing effects of the causal mask and relative positional encodings, such as the decay mask and rotary positional encoding (RoPE): while both mechanisms introduce distance-based decay within individual attention maps, their aggregate effect across multiple attention layers -- coupled with the causal mask -- leads to a trade-off between the long-term decay effects and the cumulative importance of early sequence positions. Through controlled numerical experiments, we not only validate our theoretical findings but also reproduce position biases observed in real-world LLMs. Our framework offers a principled foundation for understanding positional biases in transformers, shedding light on the complex interplay of attention mechanism components and guiding more informed architectural design.

📄 PDF Abstract BibTeX arXiv:2502.01951

Code (0)

등록된 구현이 없습니다.

Tasks

Position

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

A Multidimensional Analysis of Social Biases in Vision Transformers

2023-08-03 · ICCV 2023 1 · Jannik Brinkmann, Paul Swoboda, Christian Bartelt

The embedding spaces of image models have been shown to encode a range of social biases such as racism and sexism. Here, we investigate specific factors that contribute to the emergence of these biases in Vision Transfor…

counterfactualFairness

When does compositional structure yield compositional generalization? A kernel theory

2024-05-26 · Samuel Lippl, Kim Stachenfeld

Compositional generalization (the ability to respond correctly to novel combinations of familiar components) is thought to be a cornerstone of intelligent behavior. Compositionally structured (e.g. disentangled) represen…

MemorizationRepresentation Learning

Emergence of compositional language in communication through noisy channel

2020-06-12 · ICML Workshop LaReL 2020 7 · Łukasz Kuciński, Paweł Kołodziej, Piotr Miłoś

In this paper, we investigate how communication through a noisy channel can lead to the emergence of compositional language. Our approach is $\mbox{end-to-end}$, allows for different inductive biases on the agents’ archi…

The Inductive Bottleneck: Data-Driven Emergence of Representational Sparsity in Vision Transformers

2025-12-08 · Kanishk Awadhiya arxiv

Vision Transformers (ViTs) lack the hierarchical inductive biases inherent to Convolutional Neural Networks (CNNs), theoretically allowing them to maintain high-dimensional representations throughout all layers. However,…

Icy: A benchmark for measuring compositional inductive bias of emergent communication models

2021-09-29 · Hugh Perkins

We present a benchmark \textsc{Icy} for measuring the compositional inductive bias of models in the context of emergent communications. We devise corrupted compositional grammars that probe for limitations in the composi…

Inductive Bias