ZeroLiers: Diminishing Large Outliers in ReLU-like Activations
As the number of learnable parameters is getting bigger and bigger, overfitting is still one of the main challenges in training DNNs. Even though DNNs with billions or even a few hundred billions of parameters are proposed and used, it is still hard to determine the appropriate training set size that prevents overfitting. In this work, we propose a new activation function, called ZeroLiers, to prevent overfitting. It eliminates the need to use Dropout and leads to better generalization when training DNNs with fully connected layers. ZeroLiers can be easily implemented by replacing large outliers in ReLU-like activations with zeros. We perform an empirical evaluation of ZeroLiers' regularization effect against Dropout. Interestingly, the validation loss decreases much faster using ZeroLiers than Dropout, and the generalization performance improves. Moreover, we train several recent DNNs with fully connected layers and investigate the effect of ZeroLiers. Specifically, we find that ZeroLiers accelerates the convergence speed of both their training and validation losses.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Representation Learning and Recovery in the ReLU Model
Rectified linear units, or ReLUs, have become the preferred activation function for artificial neural networks. In this paper we consider two basic learning problems assuming that the underlying data follow a generative …
Dictionary LearningRepresentation LearningOutliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
Post-Training Quantization (PTQ) enhances the efficiency of Large Language Models (LLMs) by enabling faster operation and compatibility with more accessible hardware through reduced memory usage, at the cost of small per…
QuantizationLEARNING DISTRIBUTIONS GENERATED BY SINGLE-LAYER RELU NETWORKS IN THE PRESENCE OF ARBITRARY OUTLIERS
We consider a set of data samples such that a constant fraction of the samples are arbitrary outliers and the rest are the output samples of a single-layer neural network (NN) with rectified linear unit (ReLU) activation…
Diminishing Batch Normalization
In this paper, we propose a generalization of the Batch Normalization (BN) algorithm, diminishing batch normalization (DBN), where we update the BN parameters in a diminishing moving average way. BN is very effective in …
Uncovering Layer-Dependent Activation Sparsity Patterns in ReLU Transformers
Previous work has demonstrated that MLPs within ReLU Transformers exhibit high levels of sparsity, with many of their activations equal to zero for any given token. We build on that work to more deeply explore how token-…