paper-with-me

홈 › Papers

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory

2026-03-27 · Juno Kim, Eshaan Nichani, Denny Wu, Alberto Bietti, Jason D. Lee arxiv

Spectral optimizers such as Muon have recently shown strong empirical performance in large-scale language model training, but the source and extent of their advantage remain poorly understood. We study this question through the linear associative memory problem, a tractable model for factual recall in transformer-based models. In particular, we go beyond orthogonal embeddings and consider Gaussian inputs and outputs, which allows the number of stored associations to greatly exceed the embedding dimension. Our main result sharply characterizes the recovery rates of one step of Muon, SGD, and Newton's method on the logistic regression loss under a power law frequency distribution. We show that the storage capacity of Muon significantly exceeds that of SGD, and even matches Newton's method while only using first-order information. Moreover, Muon saturates at a larger critical batch size. We further analyze the multi-step dynamics under a thresholded gradient approximation and show that Muon achieves a substantially faster initial recovery rate than SGD, while both methods eventually converge to the information-theoretic limit at comparable speeds. Experiments on synthetic tasks validate the predicted scaling laws. Our analysis provides a quantitative understanding of the signal amplification of spectral preconditioners and lays the groundwork for establishing scaling laws across more practical language modeling tasks and optimizers.

📄 PDF Abstract BibTeX arXiv:2603.26554

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws

2026-05-20 · Nandan Kumar Jha, Brandon Reagen arxiv

Scaling laws have made language-model performance predictable from model size, data, and compute, but they typically treat the optimizer as a fixed training detail. We show that this assumption misses a fundamental axis …

Self-Organization and Spectral Mechanism of Attractor Landscapes in High-Capacity Kernel Hopfield Networks

2025-11-17 · Akira Tamamori arxiv

Kernel-based learning methods can dramatically increase the storage capacity of Hopfield networks, yet the dynamical mechanisms behind this enhancement remain poorly understood. We address this gap by combining a geometr…

Hardware-Adaptive and Superlinear-Capacity Memristor-based Associative Memory

2025-05-19 · Chengping He, Mingrui Jiang, Keyi Shan, Szu-Hao Yang 외

Brain-inspired computing aims to mimic cognitive functions like associative memory, the ability to recall complete patterns from partial cues. Memristor technology offers promising hardware for such neuromorphic systems …

FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained Control

2026-04-21 · Pingwei Sun, Yuxuan Hu, Jianchao Tan, Xue Wang 외 arxiv

Linear attention mechanisms have emerged as promising alternatives to softmax attention, offering linear-time complexity during inference. Recent advances such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA) have …

Long-Context UnderstandingComputational Efficiency

Geometric Entropy and Retrieval Phase Transitions in Continuous Thermal Dense Associative Memory

2026-04-08 · Tatiana Petrova, Evgeny Polyachenko, Radu State arxiv

We study the thermodynamic memory capacity of modern Hopfield networks (Dense Associative Memory models) with continuous states under geometric constraints, extending classical analyses of pairwise associative memory. We…