paper-with-me

홈 › Papers

Self-Attention as a Covariance Readout: A Unified View of In-Context Learning and Repetition

2026-05-11 · Haoren Xu, Guanhua Fang arxiv

Large language models (LLMs) exhibit two striking and ostensibly unrelated behaviours: in-context learning (ICL) and repetitive generation. In both, the model behaves as though it had summarised the context into a population-level statistic and discarded token-level detail. We ask whether this ``summarisation and forgetting'' can be derived from the attention mechanism itself, and answer in the affirmative. Under stationary, ergodic and elliptical inputs, the softmax attention output converges almost surely to $Θ_VΣΘ_K^{\top}Θ_Q x_t$, where $Σ$ is the input covariance; the long-context limit is therefore a linear readout of the input's second-order statistics. Two consequences follow. (i) For in-context linear regression, a single softmax head can implement one step of population gradient descent. Stacking such heads with residual connections iterates this update and implements multiple gradient descent steps. (ii) Propagated across an $L$-layer transformer, this readout drives the terminal hidden state at the parametric $1/t$ rate to a deterministic function of the current token alone, so that autoregressive generation collapses asymptotically to a first-order Markov chain whose attracting orbits furnish a structural account of repetition and mode collapse. The two phenomena thus emerge as facets of a single covariance-readout principle.

📄 PDF Abstract BibTeX arXiv:2605.10466

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Geometry-Grounded Unified 3D Perception for Autonomous Driving

2026-08-13 · Longfei Xu, Xiaohui Wang, Zehao Huang, Han Li 외 arxiv

Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pr…

3D Object DetectionAutonomous DrivingDepth Estimation

Innovation Capacity of Dynamical Learning Systems

2026-01-12 · Anthony M. Polloreno arxiv

In noisy physical reservoirs, the classical information-processing capacity $C_{\mathrm{ip}}$ quantifies how well a linear readout can realize tasks measurable from the input history, yet $C_{\mathrm{ip}}$ can be far sma…

DINO-MVR: Multi-View Readout of Frozen DINOv3 for Annotation-Efficient Medical Segmentation

2026-05-08 · Wei Jiang, Feng Liu, Nan Ye, Hongfu Sun arxiv

Adapting foundation models to medical segmentation typically requires either backbone fine-tuning or high-capacity task-specific decoders, both of which are difficult to fit reliably when annotations are scarce. We show …

Tumor Segmentation

Evaluating Representations with Readout Model Switching

2023-02-19 · Yazhe Li, Jorg Bornschein, Marcus Hutter

Although much of the success of Deep Learning builds on learning good representations, a rigorous method to evaluate their quality is lacking. In this paper, we treat the evaluation of representations as a model selectio…

modelModel Selection

Memory by Design: Probabilistic Sequence Layers

2026-05-29 · Matthew Dowling, Hyungju Jeon, Cristina Savin, Il Memming Park arxiv

We introduce the \emph{design-model framework}: a way to derive efficient recurrent sequence maps from explicit assumptions about memory. A design model writes evidence into memory by exact Bayesian filtering; a query- d…