paper-with-me

홈 › Papers

Context Staircase: Signature-Aligned Dynamics of Token Embeddings under Small Initialization

2026-08-31 · Junjie Yao, Liangkai Hang, Zhi-Qin John Xu arxiv

Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern language models learn embeddings from random initialization through gradient-based training, the dynamical mechanism by which meaningful embedding structures emerge remains unclear. In this work, we identify that the evolving embedding structures are closely related to token-conditioned label and contextual distributions, which we formalize as probability signatures. We observe a progressive learning process, which we term Context Staircase: embeddings learn the low-order statistic signatures of the data before the high-order ones. More specifically, we observe that early in training they align with the simplest, context-free signature linking a token to its label, and as training proceeds, they progressively reflect signatures involving more and more context tokens. We then analyze the gradient flow of embeddings under small initialization to explain this phenomenon, deriving embedding evolution equations for feed-forward and self-attention architectures. We further extend these observations to real language-model training. Finally, we show that these embedding structures play an important role in both task learning and the incorporation of semantic structure into the embedding space. Overall, our results provide a dynamic explanation of how data statistics and architecture jointly shape token embeddings in language models, and reveal an implicit bias in the space of data statistics: training proceeds from simpler, low-order statistical relations toward increasingly complex, context-dependent ones.

📄 PDF Abstract BibTeX arXiv:2608.30315

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Internal Flow Signatures for Self-Checking and Refinement in LLMs

2026-02-02 · Sungheon Jeong, Sanggeon Yun, Ryozo Masukawa, Wenjun Haung 외 arxiv

Large language models can generate fluent answers that are unfaithful to the provided context, while many safeguards rely on external verification or a separate judge after generation. We introduce \emph{internal flow si…

Staircase Attention for Recurrent Processing of Sequences

2021-06-08 · Da Ju, Stephen Roller, Sainbayar Sukhbaatar, Jason Weston

Attention mechanisms have become a standard tool for sequence modeling tasks, in particular by stacking self-attention layers over the entire input sequence as in the Transformer architecture. In this work we introduce a…

Language ModelingLanguage Modelling

Probability Signature: Bridging Data Semantics and Embedding Structure in Language Models

2025-09-24 · Junjie Yao, Zhi-Qin John Xu arxiv

The embedding space of language models is widely believed to capture the semantic relationships; for instance, embeddings of digits often exhibit an ordered structure that corresponds to their natural sequence. However, …

Staircase Streaming for Low-Latency Multi-Agent Inference

2025-10-06 · Junlin Wang, Jue Wang, Zhen, Xu 외 arxiv

Recent advances in large language models (LLMs) opened up new directions for leveraging the collective expertise of multiple LLMs. These methods, such as Mixture-of-Agents, typically employ additional inference steps to …

Effective Rank and the Staircase Phenomenon: New Insights into Neural Network Training Dynamics

2024-12-06 · Jiang Yang, Yuxiang Zhao, Quanhui Zhu

In recent years, deep learning, powered by neural networks, has achieved widespread success in solving high-dimensional problems, particularly those with low-dimensional feature structures. This success stems from their …

Learning Theory