paper-with-me

홈 › Papers

The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior

2026-03-30 · Laura Gomezjurado Gonzalez arxiv

Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization, but the source of that delay remains poorly understood. In encoder-decoder arithmetic models, we argue that this delay reflects limited access to already learned structure rather than failure to acquire that structure in the first place. We study one-step Collatz prediction and find that the encoder organizes parity and residue structure within the first few thousand training steps, while output accuracy remains near chance for tens of thousands more. Causal interventions support the decoder bottleneck hypothesis. Transplanting a trained encoder into a fresh model accelerates grokking by 2.75 times, while transplanting a trained decoder actively hurts. Freezing a converged encoder and retraining only the decoder eliminates the plateau entirely and yields 97.6% accuracy, compared to 86.1% for joint training. What makes the decoder's job harder or easier depends on numeral representation. Across 15 bases, those whose factorization aligns with the Collatz map's arithmetic (e.g., base 24) reach 99.8% accuracy, while binary fails completely because its representations collapse and never recover. The choice of base acts as an inductive bias that controls how much local digit structure the decoder can exploit, producing large differences in learnability from the same underlying task.

📄 PDF Abstract BibTeX arXiv:2604.13082

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Early-Warning Signals of Grokking via Loss-Landscape Geometry

2026-02-19 · Yongzhong Xu arxiv

Grokking -- the abrupt transition from memorization to generalization after prolonged training -- has been linked to confinement on low-dimensional execution manifolds in modular arithmetic. Whether this mechanism extend…

A Bayesian Perspective on the Role of Epistemic Uncertainty for Delayed Generalization in In-Context Learning

2026-04-14 · Abdessamed Qchohi, Simone Rossi arxiv

In-context learning enables transformers to adapt to new tasks from a few examples at inference time, while grokking highlights that this generalization can emerge abruptly only after prolonged training. We study task ge…

Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization

2026-06-30 · Srijan Tiwari, Aditya Chauhan, Manjot Singh arxiv

Why do neural networks memorize algorithmic training data long before they generalize? We present a geometric case study demonstrating that, on tasks where generalization requires discovering structured low-dimensional c…

Breaking Data Symmetry is Needed For Generalization in Feature Learning Kernels

2026-03-31 · Marcel Tomàs Bernal, Neil Rohit Mallinar, Mikhail Belkin arxiv

Grokking occurs when a model achieves high training accuracy but generalization to unseen test points happens long after that. This phenomenon was initially observed on a class of algebraic problems, such as learning mod…

Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker Model

2025-04-17 · Zhiwei Xu, Zhiyu Ni, Yixin Wang, Wei Hu

''Grokking'' is a phenomenon where a neural network first memorizes training data and generalizes poorly, but then suddenly transitions to near-perfect generalization after prolonged training. While intriguing, this dela…