paper-with-me

Papers

Constrained belief updates explain geometric structures in transformer representations

2025-02-04 · Mateusz Piotrowski, Paul M. Riechers, Daniel Filan, Adam S. Shai

What computational structures emerge in transformers trained on next-token prediction? In this work, we provide evidence that transformers implement constrained Bayesian belief updating -- a parallelized version of partial Bayesian inference shaped by architectural constraints. To do this, we integrate the model-agnostic theory of optimal prediction with mechanistic interpretability to analyze transformers trained on a tractable family of hidden Markov models that generate rich geometric patterns in neural activations. We find that attention heads carry out an algorithm with a natural interpretation in the probability simplex, and create representations with distinctive geometric structure. We show how both the algorithmic behavior and the underlying geometry of these representations can be theoretically predicted in detail -- including the attention pattern, OV-vectors, and embedding vectors -- by modifying the equations for optimal future token predictions to account for the architectural constraints of attention. Our approach provides a principled lens on how gradient descent resolves the tension between optimal prediction and architectural design.

📄 PDF Abstract BibTeX arXiv:2502.01954

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian InferencePrediction

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Confidence is Not Competence

2025-10-24 · Debdeep Sanyal, Manya Pandey, Dhruv Kumar, Saurabh Deshpande 외 arxiv

Large language models (LLMs) often exhibit a puzzling disconnect between their asserted confidence and actual problem-solving competence. We offer a mechanistic account of this decoupling by analyzing the geometry of int…

Chance-Constrained Active Inference

2021-02-17 · Thijs van de Laar, Ismail Senoz, Ayça Özçelikkale, Henk Wymeersch

Active Inference (ActInf) is an emerging theory that explains perception and action in biological agents, in terms of minimizing a free energy bound on Bayesian surprise. Goal-directed behavior is elicited by introducing…

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space

2026-05-12 · Eric Bigelow, Raphaël Sarfati, Daniel Wurgaft, Owen Lewis 외 arxiv

Large Language Models (LLMs) update their behavior in context, which can be viewed as a form of Bayesian inference. However, the structure of the latent hypothesis space over which this inference operates remains unclear…

Bayesian Inference

Constrained Diffusion for Protein Design with Hard Structural Constraints

2025-10-01 · Jacob K. Christopher, Austin Seamann, Jingyi Cui, Sagar Khare 외 arxiv

Diffusion models offer a powerful means of capturing the manifold of realistic protein structures, enabling rapid design for protein engineering tasks. However, existing approaches observe critical failure modes when pre…

Protein Design

Geometric and Dynamic Scaling in Deep Transformers

2026-01-03 · Haoran Su, Chenyu You arxiv

Despite their empirical success, pushing Transformer architectures to extreme depth often leads to a paradoxical failure: representations become increasingly redundant, lose rank, and ultimately collapse. Existing explan…

Representation Learning