paper-with-me

홈 › Papers

A Geometric Perspective on Next-Token Prediction in Large Language Models: Three Emerging Phases

2026-05-09 · Gianfranco Lombardo, Giuseppe Trimigno, Stefano Cagnoni arxiv

We investigate the geometry of predictive information across the layers of large language models (LLMs). We repurpose representation lenses-learned affine maps trained to predict the next token from intermediate residual streams-as geometric diagnostic tools. Rather than asking what the model predicts at each layer, we ask where predictive information resides and how it evolves across depth. We define at each layer a predictive readout subspace as the dominant k-dimensional singular subspace of such a map on the d-dimensional residual stream (where k is a resolution parameter), and track its trajectory on the Grassmann manifold as a similarity profile across layers. The profile is well described by unimodal distributions exhibiting a rise, near-plateau, and descent; varying k from 1% to 50% of d traces a Pareto frontier between visibility and energy retention, yet the same structure emerges at all scales. Across eight models from two families (Qwen2.5 and OLMo2, 1B-32B), we identify three geometric phases. Updates are approximately orthogonal to the residual stream throughout; what distinguishes the phases is their effect on the effective rank, which expands, stabilizes, and concentrates. In the first, Seeding Multiplexing, feed-forward memories and attention layers seed a candidate set in superposition in family-specific proportions, with the final token rising as leading candidate from 20% to 35% of positions across this phase. In the second, Hoisting Overriding, updates override existing subspaces to concentrate the candidate distribution without expanding the rank. In the third, Focal Convergence, high-energy low-rank updates write the winner into a form aligned with the unembedding direction. Phases 1 and 3 grow slowly with model depth, while Phase 2 expands linearly. The additional capacity of deeper LLMs is largely absorbed by candidate disambiguation.

📄 PDF Abstract BibTeX arXiv:2605.09011

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emergent Causal-Geometric Dynamics Across Depth in Large Language Models

2026-02-04 · Shahar Haim, Daniel C McNamee arxiv

Geometric analyses of large language model (LLM) representations reveal structured variation across depth but remain fundamentally correlational with respect to token prediction formation. Meanwhile, causal interventions…

Rethinking Point Clouds as Sequences: A Causal Next-Token Predictive Learning Framework

2026-05-17 · Yumeng Yao, Jingzhi Dong, Haowen Gu, Tao Chen 외 arxiv

With the rapid progress of multimodal foundation models and predictive pre-training, an important open question is how to equip 3D point clouds with a pre-training paradigm that is better aligned with next-token and next…

Self-Supervised LearningPoint Clouds

Representational Curvature Modulates Behavioral Uncertainty in Large Language Models

2026-04-27 · Jack King, Evelina Fedorenko, Eghbal A. Hosseini arxiv

In autoregressive large language models (LLMs), temporal straightening offers an account of how the next-token prediction objective shapes representations. Models learn to progressively straighten the representational tr…

The Geometry of Tokens in Internal Representations of Large Language Models

2025-01-17 · Karthik Viswanathan, Yuri Gardinazzi, Giada Panerai, Alberto Cazzaniga 외

We investigate the relationship between the geometry of token embeddings and their role in the next token prediction within transformer models. An important aspect of this connection uses the notion of empirical measure,…

A Law of Next-Token Prediction in Large Language Models

2024-08-24 · Hangfeng He, Weijie J. Su

Large language models (LLMs) have been widely employed across various application domains, yet their black-box nature poses significant challenges to understanding how these models process input data internally to make p…

Mamba