Understanding Why Neural Networks Generalize Well Through GSNR of Parameters
As deep neural networks (DNNs) achieve tremendous success across many application domains, researchers tried to explore in many aspects on why they generalize well. In this paper, we provide a novel perspective on these issues using the gradient signal to noise ratio (GSNR) of parameters during training process of DNNs. The GSNR of a parameter is defined as the ratio between its gradient's squared mean and variance, over the data distribution. Based on several approximations, we establish a quantitative relationship between model parameters' GSNR and the generalization gap. This relationship indicates that larger GSNR during training process leads to better generalization performance. Moreover, we show that, different from that of shallow models (e.g. logistic regression, support vector machines), the gradient descent optimization dynamics of DNNs naturally produces large GSNR during training, which is probably the key to DNNs' remarkable generalization ability.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Unleashing the Power of Gradient Signal-to-Noise Ratio for Zero-Shot NAS
Neural Architecture Search (NAS) aims to automatically find optimal neural network architectures in an efficient way. Zero-Shot NAS is a promising technique that leverages proxies to predict the accuracy of candidate…
Neural Architecture SearchDomain Generalization Guided by Gradient Signal to Noise Ratio of Parameters
Overfitting to the source domain is a common issue in gradient-based training of deep neural networks. To compensate for the over-parameterized models, numerous regularization techniques have been introduced such as thos…
Domain GeneralizationFace Anti-SpoofingMeta-LearningFast WDM provisioning with minimal probing: the first field experiments for DC exchanges
We propose an approach to estimate the end-to-end GSNR accurately in a short time when a data center interconnect (DCI) network operator receives a service request from users, not by measuring the GSNR at the operational…
GSNR: Graph Smooth Null-Space Representation for Inverse Problems
Inverse problems in imaging are ill-posed, leading to infinitely many solutions consistent with the measurements due to the non-trivial null-space of the sensing matrix. Common image priors promote solutions on the gener…
Image Super-ResolutionImage DeblurringLaunch Power Optimization for Dynamic Elastic Optical Networks over C+L Bands
We propose an algorithm for calculating the optimum launch power over the entire C+L bands by maximizing the cumulative link GSNR of a channel plan built upon multiple modulation formats, with application to dynamic EONs…