paper-with-me

Papers

Flatness is Necessary, Neural Collapse is Not: Rethinking Generalization via Grokking

2025-09-22 · Ting Han, Linara Adilova, Henning Petzka, Jens Kleesiek, Michael Kamp arxiv

Neural collapse, i.e., the emergence of highly symmetric, class-wise clustered representations, is frequently observed in deep networks and is often assumed to reflect or enable generalization. In parallel, flatness of the loss landscape has been theoretically and empirically linked to generalization. Yet, the causal role of either phenomenon remains unclear: Are they prerequisites for generalization, or merely by-products of training dynamics? We disentangle these questions using grokking, a training regime in which memorization precedes generalization, allowing us to temporally separate generalization from training dynamics and we find that while both neural collapse and relative flatness emerge near the onset of generalization, only flatness consistently predicts it. Models encouraged to collapse or prevented from collapsing generalize equally well, whereas models regularized away from flat solutions exhibit delayed generalization, resembling grokking, even in architectures and datasets where it does not typically occur. Furthermore, we show theoretically that neural collapse leads to relative flatness under classical assumptions, explaining their empirical co-occurrence. Our results support the view that relative flatness is a potentially necessary and more fundamental property for generalization, and demonstrate how grokking can serve as a powerful probe for isolating its geometric underpinnings.

📄 PDF Abstract BibTeX arXiv:2509.17738

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Late-Stage Generalization Collapse in Grokking: Detecting anti-grokking with Weightwatcher

2026-02-02 · Hari K Prakash, Charles H Martin arxiv

\emph{Memorization} in neural networks lacks a precise operational definition and is often inferred from the grokking regime, where training accuracy saturates while test accuracy remains very low. We identify a previous…

Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory

2025-06-04 · Hari K. Prakash, Charles H. Martin

We study the well-known grokking phenomena in neural networks (NNs) using a 3-layer MLP trained on 1 k-sample subset of MNIST, with and without weight decay, and discover a novel third phase -- \emph{anti-grokking} -- th…

Grokking at the Edge of Numerical Stability

2025-01-08 · Lucas Prieto, Melih Barsbey, Pedro A. M. Mediano, Tolga Birdal

Grokking, the sudden generalization that occurs after prolonged overfitting, is a surprising phenomenon challenging our understanding of deep learning. Although significant progress has been made in understanding grokkin…

Canalization Before Generalization: Grokking as a Dynamical Probe

2026-08-26 · Yiming Lin arxiv

For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a win…

Grokking From Abstraction to Intelligence

2026-03-31 · Junjie Zhang, Zhen Shen, Gang Xiong, Xisong Dong arxiv

Grokking in modular arithmetic has established itself as the quintessential fruit fly experiment, serving as a critical domain for investigating the mechanistic origins of model generalization. Despite its significance, …