paper-with-me

홈 › Papers

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

2025-08-18 · Aditya Varre, Gizem Yüce, Nicolas Flammarion arxiv

Motivated by empirical observations of prolonged plateaus and stage-wise progression during training, we investigate the loss landscape of transformer models trained on in-context next-token prediction tasks. In particular, we focus on learning in-context $n$-gram language models under cross-entropy loss, and establish a sufficient condition for parameter configurations to be stationary points. We then construct a set of parameter configurations for a simplified transformer model that represent $k$-gram estimators (for $k \leq n$), and show that the gradient of the population loss at these solutions vanishes in the limit of infinite sequence length and parameter norm. This reveals a key property of the loss landscape: {sub-$n$-grams are near-stationary points of the population cross-entropy loss}, offering theoretical insight into widely observed phenomena such as stage-wise learning dynamics and emergent phase transitions. These insights are further supported by numerical experiments that illustrate the learning dynamics of $n$-grams, characterized by discrete transitions between near-stationary solutions.

📄 PDF Abstract BibTeX arXiv:2508.12837

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Incremental Path-Following Splitting Method for Linearly Constrained Nonconvex Nonsmooth Programs

2018-01-30 · Linbo Qiao, Wei Liu, Steven Hoi

The stationary point of Problem 2 is NOT the stationary point of Problem 1. We are sorry and we are working on fixing this error.

Learning Transformer Programs

2023-06-01 · NeurIPS 2023 11 · Dan Friedman, Alexander Wettig, Danqi Chen

Recent research in mechanistic interpretability has attempted to reverse-engineer Transformer models by carefully inspecting network weights and activations. However, these approaches require considerable manual effort a…

In-Context LearningInterpretable Machine Learningnamed-entity-recognitionNamed Entity Recognition+2

Universal Length Generalization with Turing Programs

2024-07-03 · Kaiying Hou, David Brandfonbrener, Sham Kakade, Samy Jelassi 외

Length generalization refers to the ability to extrapolate from short training sequences to long test sequences and is a challenge for current large language models. While prior work has proposed some architecture or dat…

Looped Transformers as Programmable Computers

2023-01-30 · Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee 외

We present a framework for using transformer networks as universal computers by programming them with specific weights and placing them in a loop. Our input sequence acts as a punchcard, consisting of instructions and me…

In-Context Learning

Discovering Interpretable Algorithms by Decompiling Transformers to RASP

2026-02-09 · Xinting Huang, Aleksandra Bakalova, Satwik Bhattamishra, William Merrill 외 arxiv

Recent work has shown that the computations of Transformers can be simulated in the RASP family of programming languages. These findings have enabled improved understanding of the expressive capacity and generalization a…