paper-with-me

홈 › Papers

Inductive Bias of Gradient Descent for Weight Normalized Smooth Homogeneous Neural Nets

2020-10-24 · Depen Morwani, Harish G. Ramaswamy

We analyze the inductive bias of gradient descent for weight normalized smooth homogeneous neural nets, when trained on exponential or cross-entropy loss. We analyse both standard weight normalization (SWN) and exponential weight normalization (EWN), and show that the gradient flow path with EWN is equivalent to gradient flow on standard networks with an adaptive learning rate. We extend these results to gradient descent, and establish asymptotic relations between weights and gradients for both SWN and EWN. We also show that EWN causes weights to be updated in a way that prefers asymptotic relative sparsity. For EWN, we provide a finite-time convergence rate of the loss with gradient flow and a tight asymptotic convergence rate with gradient descent. We demonstrate our results for SWN and EWN on synthetic data sets. Experimental results on simple datasets support our claim on sparse EWN solutions, even with SGD. This demonstrates its potential applications in learning neural networks amenable to pruning.

📄 PDF Abstract BibTeX arXiv:2010.12909

Code (1)

DepenM/Exp-WN 공식 구현 pytorch

Tasks

Inductive Bias

Methods 이 논문이 사용한 방법론

Weight Normalization Weight Normalization is a normalization method for training neural networks. It is inspired by batch normalization,…
SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

The Implicit Bias of Gradient Descent on Generalized Gated Linear Networks

2022-02-05 · Samuel Lippl, L. F. Abbott, SueYeon Chung

Understanding the asymptotic behavior of gradient-descent training of deep neural networks is essential for revealing inductive biases and improving network performance. We derive the infinite-time training limit of a ma…

Inductive Bias

The late-stage training dynamics of (stochastic) subgradient descent on homogeneous neural networks

2025-02-08 · Sholom Schechtman, Nicolas Schreuder

We analyze the implicit bias of constant step stochastic subgradient descent (SGD). We consider the setting of binary classification with homogeneous neural networks - a large class of deep neural networks with ReLU-type…

Binary Classification

Inductive Bias of Multi-Channel Linear Convolutional Networks with Bounded Weight Norm

2021-02-24 · Meena Jagadeesan, Ilya Razenshteyn, Suriya Gunasekar

We provide a function space characterization of the inductive bias resulting from minimizing the $\ell_2$ norm of the weights in multi-channel convolutional neural networks with linear activations and empirically test ou…

Inductive Bias

Neural Redshift: Random Networks are not Random Functions

2024-03-04 · CVPR 2024 1 · Damien Teney, Armand Nicolicioiu, Valentin Hartmann, Ehsan Abbasnejad

Our understanding of the generalization capabilities of neural networks (NNs) is still incomplete. Prevailing explanations are based on implicit biases of gradient descent (GD) but they cannot account for the capabilitie…

Depth Without the Magic: Inductive Bias of Natural Gradient Descent

2021-11-22 · Anna Kerekes, Anna Mészáros, Ferenc Huszár

In gradient descent, changing how we parametrize the model can lead to drastically different optimization trajectories, giving rise to a surprising range of meaningful inductive biases: identifying sparse classifiers or …

Inductive Bias