paper-with-me

홈 › Papers

Variance Control via Weight Rescaling in LLM Pre-training

2025-03-21 · Louis Owen, Abhay Kumar, Nilabhra Roy Chowdhury, Fabian Güra

The outcome of Large Language Model (LLM) pre-training strongly depends on weight initialization and variance control strategies. Although the importance of initial variance control has been well documented in neural networks in general, the literature on initialization and management of its growth during LLM pre-training, specifically, is somewhat sparse. In this paper, we introduce the Layer Index Rescaling (LIR) weight initialization scheme, and the Target Variance Rescaling (TVR) variance control strategy. Experiments on a 1B parameter LLaMA model demonstrate that better variance management using these techniques yields substantial improvements in downstream task performance (up to 4.6% on common pre-training benchmarks) and reduces extreme activation values, thus mitigating challenges associated with quantization and low-precision training. Our code is available at: https://github.com/bluorion-com/weight_rescaling.

📄 PDF Abstract BibTeX arXiv:2503.17500

Code (1)

bluorion-com/weight_rescaling 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelManagementQuantization

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

Non-Vacuous Generalization Bounds: Can Rescaling Invariances Help?

2025-09-30 · Damien Rouchouse, Antoine Gonon, Rémi Gribonval, Benjamin Guedj arxiv

A central challenge in understanding generalization is to obtain non-vacuous guarantees that go beyond worst-case complexity over data or weight space. Among existing approaches, PAC-Bayes bounds stand out as they can pr…

Weight Rescaling: Effective and Robust Regularization for Deep Neural Networks with Batch Normalization

2021-02-06 · Ziquan Liu, Yufei Cui, Jia Wan, Yu Mao 외

Weight decay is often used to ensure good generalization in the training practice of deep neural networks with batch normalization (BN-DNNs), where some convolution layers are invariant to weight rescaling due to the nor…

Crowd Countingimage-classificationImage Classificationobject-detection+2

How to Transform Kernels for Scale-Convolutions

2021-07-26 · ICCVW 2021 7 · Ivan Sosnovik, Artem Moskalev, Arnold Smeulders

Scale is often seen as a given, disturbing factor in many vision tasks. When doing so it is one of the factors why we need more data during learning. In recent work scale equivariance was added to convolutional neural ne…

Rescaling CNN through Learnable Repetition of Network Parameters

2021-01-14 · Arnav Chavan, Udbhav Bamba, Rishabh Tiwari, Deepak Gupta

Deeper and wider CNNs are known to provide improved performance for deep learning tasks. However, most such networks have poor performance gain per parameter increase. In this paper, we investigate whether the gain obser…

An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies

2026-07-22 · Jiachun Li arxiv

Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood. This preliminary report pr…