paper-with-me

Papers

Critical Data Size of Language Models from a Grokking Perspective

2024-01-19 · Xuekai Zhu, Yao Fu, BoWen Zhou, Zhouhan Lin

We explore the critical data size in language models, a threshold that marks a fundamental shift from quick memorization to slow generalization. We formalize the phase transition under the grokking configuration into the Data Efficiency Hypothesis and identify data insufficiency, sufficiency, and surplus regimes in language models training dynamics. We develop a grokking configuration to reproduce grokking on simplistic language models stably by rescaling initialization and weight decay. We show that generalization occurs only when language models reach a critical size. We analyze grokking across sample-wise and model-wise, verifying the proposed data efficiency hypothesis. Our experiments reveal smoother phase transitions occurring at the critical dataset size for language datasets. As the model size increases, this critical point also becomes larger, indicating that larger models require more data. Our results deepen the understanding of language model training, offering a novel perspective on the role of data in the learning mechanism of language models.

📄 PDF Abstract BibTeX arXiv:2401.10463

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMemorization

Similar Papers 제목 키워드 기반

Grokking as a Falsifiable Finite-Size Transition

2026-03-25 · Yuda Bi, Chenyu Zhang, Qiheng Wang, Vince D Calhoun arxiv

Grokking -- the delayed onset of generalization after early memorization -- is often described with phase-transition language, but that claim has lacked falsifiable finite-size inputs. Here we supply those inputs by trea…

Unified View of Grokking, Double Descent and Emergent Abilities: A Perspective from Circuits Competition

2024-02-23 · Yufei Huang, Shengding Hu, Xu Han, Zhiyuan Liu 외

Recent studies have uncovered intriguing phenomena in deep learning, such as grokking, double descent, and emergent abilities in large language models, which challenge human intuition and are crucial for a deeper underst…

MemorizationMulti-Task Learning

Language Models "Grok" to Copy

2024-09-14 · Ang Lv, Ruobing Xie, Xingwu Sun, Zhanhui Kang 외

We examine the pre-training dynamics of language models, focusing on their ability to copy text from preceding context--a fundamental skill for various LLM applications, including in-context learning (ICL) and retrieval-…

In-Context LearningLanguage ModellingRAGRetrieval-augmented Generation

Grokking as Dimensional Phase Transition in Neural Networks

2026-04-06 · Ping Wang arxiv

Neural network grokking -- the abrupt memorization-to-generalization transition -- challenges our understanding of learning dynamics. Through finite-size scaling of gradient avalanche dynamics across eight model scales, …

Grokking phase transitions in learning local rules with gradient descent

2022-10-26 · Bojan Žunkovič, Enej Ilievski

We discuss two solvable grokking (generalisation beyond overfitting) models in a rule learning scenario. We show that grokking is a phase transition and find exact analytic expressions for the critical exponents, grokkin…

Learning Theory