paper-with-me

홈 › Papers

Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy

2025-10-05 · Karthik Viswanathan, Sang Eon Park arxiv

We introduce a cumulant-expansion framework for quantifying how large language models (LLMs) internalize higher-order statistical structure during next-token prediction. By treating the softmax entropy of each layer's logit distribution as a perturbation around its "center" distribution, we derive closed-form cumulant observables that isolate successively higher-order correlations. Empirically, we track these cumulants in GPT-2 and Pythia models on Pile-10K prompts. (i) Structured prompts exhibit a characteristic rise-and-plateau profile across layers, whereas token-shuffled prompts remain flat, revealing the dependence of the cumulant profile on meaningful context. (ii) During training, all cumulants increase monotonically before saturating, directly visualizing the model's progression from capturing variance to learning skew, kurtosis, and higher-order statistical structures. (iii) Mathematical prompts show distinct cumulant signatures compared to general text, quantifying how models employ fundamentally different processing mechanisms for mathematical versus linguistic content. Together, these results establish cumulant analysis as a lightweight, mathematically grounded probe of feature-learning dynamics in high-dimensional neural networks.

📄 PDF Abstract BibTeX arXiv:2510.04285

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing

2026-03-18 · Raghavv Goel, Mukul Gagrani, Mingu Lee, Chris Lott arxiv

Large Language Models (LLMs) possess latent multi-token prediction (MTP) abilities despite being trained only for next-token generation. We introduce ESP (Embedding-Space Probing), a simple and training-free MTP method t…

Large Language Models Develop Belief State Geometry In-Context

2026-09-15 · Daniel Balcells, Andrew Jun Lee, Chirag Rastogi, Paul M. Riechers 외 arxiv

Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in…

The Geometry of Tokens in Internal Representations of Large Language Models

2025-01-17 · Karthik Viswanathan, Yuri Gardinazzi, Giada Panerai, Alberto Cazzaniga 외

We investigate the relationship between the geometry of token embeddings and their role in the next token prediction within transformer models. An important aspect of this connection uses the notion of empirical measure,…

Inference Time Causal Probing in LLMs

2026-05-08 · Sadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash, Matthias Grossglauser arxiv

Causal probing methods aim to test and control how internal representations influence the behavior of generative models. In causal probing, an intervention modifies hidden states so that a property takes on a different v…

Transformers represent belief state geometry in their residual stream

2024-05-24 · Adam S. Shai, Sarah E. Marzen, Lucas Teixeira, Alexander Gietelink Oldenziel 외

What computational structure are we building into large language models when we train them on next-token prediction? Here, we present evidence that this structure is given by the meta-dynamics of belief updating over hid…

Prediction