paper-with-me

홈 › Papers

Uniform Scaling Limits in AdamW-Trained Transformers

2026-05-11 · William Gibson, Christoph Reisinger arxiv

We study the large-depth limit of transformers trained with AdamW, by modelling the hidden-state dynamics as an interacting particle system (IPS) coupled through the attention mechanism. Under appropriate scaling of the attention heads, we prove that the joint dynamics of the hidden states and backpropagated variables converge in $L^2$, uniformly over the initial condition, to the solution of a forward--backward system of ODEs at rate $\mathcal O(L^{-1}+L^{-1/3}H^{-1/2})$. Here, $L$ and $H$ denote the depth and number of heads of the transformer, respectively. The limiting system of ODEs can be identified with a McKean--Vlasov ODE (MVODE) when the attention heads do not incorporate causal masking. By using the flow maps associated with this MVODE and applying concentration of measure techniques, we obtain bounds on the difference between the discrete and continuous models that are uniform over compact sets of initial conditions. As this is achieved without resorting to a covering argument, the constants in our bounds are independent of the number of tokens. Furthermore, under a suitable adaptation to AdamW, the bounds become independent of the token embedding dimension.

📄 PDF Abstract BibTeX arXiv:2605.11059

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How to set AdamW's weight decay as you scale model and dataset size

2024-05-22 · Xi Wang, Laurence Aitchison

The scaling of the optimal AdamW weight decay hyperparameter with model and dataset size is critical as we seek to build larger models, but is poorly understood. We show that weights learned by AdamW can be understood as…

No More Adam: Learning Rate Scaling at Initialization is All You Need

2024-12-16 · Minghao Xu, Lichuan Xiang, Xu Cai, Hongkai Wen

In this work, we question the necessity of adaptive gradient methods for training deep neural networks. SGD-SaI is a simple yet effective enhancement to stochastic gradient descent with momentum (SGDM). SGD-SaI performs …

All

When Does Muon Help Agentic Reinforcement Learning?

2026-07-17 · Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou 외 arxiv

Muon is competitive with AdamW in large-scale pre-training, but its operating regime in reinforcement-learning post-training remains unclear. We map this regime on ALFWorld, a sparse-reward agentic benchmark, using three…

Reinforcement Learning

Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs

2026-02-07 · Sagnik Mukherjee, Lifan Yuan, Pavan Jayasinha, Dilek Hakkani-Tür 외 arxiv

Reinforcement learning (RL), particularly RL from verifiable reward (RLVR), has become a crucial phase of training large language models (LLMs) and a key focus of current scaling efforts. However, optimization practices …

Reinforcement Learning

Scaling Muon for Diffusion Transformers

2026-08-21 · Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu 외 arxiv

The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion Transformers (DiTs) remain unclear. We …