paper-with-me

홈 › Papers

Understanding the Disharmony between Weight Normalization Family and Weight Decay: $ε-$shifted $L_2$ Regularizer

2019-11-14 · Li Xiang, Chen Shuo, Xia Yan, Yang Jian

The merits of fast convergence and potentially better performance of the weight normalization family have drawn increasing attention in recent years. These methods use standardization or normalization that changes the weight $\boldsymbol{W}$ to $\boldsymbol{W}'$, which makes $\boldsymbol{W}'$ independent to the magnitude of $\boldsymbol{W}$. Surprisingly, $\boldsymbol{W}$ must be decayed during gradient descent, otherwise we will observe a severe under-fitting problem, which is very counter-intuitive since weight decay is widely known to prevent deep networks from over-fitting. In this paper, we \emph{theoretically} prove that the weight decay term $\frac{1}{2}\lambda||{\boldsymbol{W}}||^2$ merely modulates the effective learning rate for improving objective optimization, and has no influence on generalization when the weight normalization family is compositely employed. Furthermore, we also expose several critical problems when introducing weight decay term to weight normalization family, including the missing of global minimum and training instability. To address these problems, we propose an $\epsilon-$shifted $L_2$ regularizer, which shifts the $L_2$ objective by a positive constant $\epsilon$. Such a simple operation can theoretically guarantee the existence of global minimum, while preventing the network weights from being too small and thus avoiding gradient float overflow. It significantly improves the training stability and can achieve slightly better performance in our practice. The effectiveness of $\epsilon-$shifted $L_2$ regularizer is comprehensively validated on the ImageNet, CIFAR-100, and COCO datasets. Our codes and pretrained models will be released in https://github.com/implus/PytorchInsight.

📄 PDF Abstract BibTeX arXiv:1911.05920

Code (1)

implus/PytorchInsight 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Weight Decay 설명 없음
Weight Normalization Weight Normalization is a normalization method for training neural networks. It is inspired by batch normalization,…

Similar Papers 제목 키워드 기반

Understanding the Disharmony between Dropout and Batch Normalization by Variance Shift

2018-01-16 · CVPR 2019 6 · Xiang Li, Shuo Chen, Xiaolin Hu, Jian Yang

This paper first answers the question "why do the two most powerful techniques Dropout and Batch Normalization (BN) often lead to a worse performance when they are combined together?" in both theoretical and statistical …

The Disharmony between BN and ReLU Causes Gradient Explosion, but is Offset by the Correlation between Activations

2023-04-23 · Inyoung Paik, Jaesik Choi

Deep neural networks, which employ batch normalization and ReLU-like activation functions, suffer from instability in the early stages of training due to the high gradient induced by temporal gradient explosion. In this …

Understanding Weight Normalized Deep Neural Networks with Rectified Linear Units

2018-10-03 · NeurIPS 2018 12 · Yixi Xu, Xiao Wang

This paper presents a general framework for norm-based capacity control for $L_{p,q}$ weight normalized deep neural networks. We establish the upper bound on the Rademacher complexities of this family. With an $L_{p,q}$ …

Consonant harmony, disharmony, memory and time scales

2021-02-01 · SCiL 2021 2 · Adamantios Gafos

MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration

2026-03-30 · Da Chang, Qiankun Shi, Lvgang Zhang, Yu Li 외 arxiv

Orthogonalized-update optimizers such as Muon improve training of matrix-valued parameters, but existing extensions typically either rescale updates after orthogonalization or use heavier whitening-based preconditioners …