paper-with-me

Papers

Canalization Before Generalization: Grokking as a Dynamical Probe

2026-08-26 · Yiming Lin arxiv

For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight-decay (WD) pulses across this plateau and measure how they shift later generalization time. Across three grokking tasks, these shifts are unordered early in the plateau but later form a stable dose ordering, with stronger WD increases leading to earlier generalization and stronger WD decreases leading to later generalization. This ordering emerges before visible generalization in all three tasks. Meanwhile, test-loss barriers between perturbed and baseline generalization checkpoints collapse toward zero while the ordered timing effects persist. We call this combination of increasingly constrained solution selection and persistent dose-ordered timing sensitivity the canalization of function selection.

📄 PDF Abstract BibTeX arXiv:2608.25813

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dimensional Criticality at Grokking Across MLPs and Transformers

2026-04-06 · Ping Wang arxiv

Abrupt transitions between distinct dynamical regimes are a hallmark of complex systems. Grokking in deep neural networks provides a striking example -- an abrupt transition from memorization to generalization long after…

The Lifecycle of the Spectral Edge: From Gradient Learning to Weight-Decay Compression

2026-04-08 · Yongzhong Xu arxiv

We decompose the spectral edge -- the dominant direction of the Gram matrix of parameter updates -- into its gradient and weight-decay components during grokking in two sequence tasks (Dyck-1 and SCAN). We find a sharp t…

Understanding Grokking Through A Robustness Viewpoint

2023-11-11 · Zhiquan Tan, Weiran Huang

Recently, an interesting phenomenon called grokking has gained much attention, where generalization occurs long after the models have initially overfitted the training data. We try to understand this seemingly strange ph…

Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking

2025-09-22 · Ting Han, Linara Adilova, Henning Petzka, Jens Kleesiek 외 arxiv

Neural collapse, i.e., the emergence of highly symmetric, class-wise clustered representations, is frequently observed in deep networks and is often assumed to reflect or enable generalization. In parallel, flatness of t…

Low-Dimensional and Transversely Curved Optimization Dynamics in Grokking

2026-02-18 · Yongzhong Xu arxiv

Grokking -- the delayed transition from memorization to generalization in small algorithmic tasks -- remains poorly understood. We present a geometric analysis of optimization dynamics in transformers trained on modular …