paper-with-me

홈 › Papers

MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration

2026-02-02 · Lianhai Ren, Yucheng Ding, Xiao Liu, Qianxiao Li, Peng Cheng, Yeyun Gong arxiv

Training instability remains a critical challenge in large language model (LLM) pretraining, often manifesting as sudden gradient explosions that waste significant computational resources. We study training failures in a 5M-parameter NanoGPT model scaled via $μ$P, identifying two key phenomena preceding collapse: (1) rapid decline in weight matrix stable rank (ratio of squared Frobenius norm to squared spectral norm), and (2) increasing alignment between adjacent layer Jacobians. We prove theoretically that these two conditions jointly cause exponential gradient norm growth with network depth. To break this instability mechanism, we propose MSign, a new optimizer that periodically applies matrix sign operations to restore stable rank. Experiments on models from 5M to 3B parameters demonstrate that MSign effectively prevents training failures with a computational overhead of less than 7.0%.

📄 PDF Abstract BibTeX arXiv:2602.01734

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multiscale Invertible Generative Networks for High-Dimensional Bayesian Inference

2021-05-12 · Shumao Zhang, Pengchuan Zhang, Thomas Y. Hou

We propose a Multiscale Invertible Generative Network (MsIGN) and associated training algorithm that leverages multiscale structure to solve high-dimensional Bayesian inference. To address the curse of dimensionality, Ms…

Bayesian InferenceImage GenerationVocal Bursts Intensity Prediction

Demystifying Manifold Constraints in LLM Pre-training

2026-05-06 · Kang An, Jiaxiang Li, Donald Goldfarb, Shiqian Ma arxiv

The empirical success of large language model (LLM) pre-training relies heavily on heuristic stabilization techniques, such as explicit normalization layers and weight decay. While recent constrained optimization approac…

Learning to Learn with Smooth Regularization

2021-01-01 · Yuanhao Xiong, Cho-Jui Hsieh

Recent decades have witnessed great prosperity of deep learning in tackling various problems such as classification and decision making. The rapid development stimulates a novel framework, Learning-to-Learn (L2L), in whi…

Decision MakingFew-Shot Learning

REG: A Regularization Optimizer for Robust Training Dynamics

2025-10-04 · Zehua Liu, Han Wu, Xiaojin Fu, Shuqi Liu 외 arxiv

Optimizers are crucial for the efficient training of Large Language Models (LLMs). While AdamW is the de facto standard, recent structure-aware optimizers like Muon have emerged, which regularize gradient updates by oper…

CAME: Confidence-guided Adaptive Memory Efficient Optimization

2023-07-05 · Yang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang 외

Adaptive gradient methods, such as Adam and LAMB, have demonstrated excellent performance in the training of large language models. Nevertheless, the need for adaptivity requires maintaining second-moment estimates of th…