paper-with-me

홈 › Papers

Experiments with Rich Regime Training for Deep Learning

2021-02-26 · Xinyan Li, Arindam Banerjee

In spite of advances in understanding lazy training, recent work attributes the practical success of deep learning to the rich regime with complex inductive bias. In this paper, we study rich regime training empirically with benchmark datasets, and find that while most parameters are lazy, there is always a small number of active parameters which change quite a bit during training. We show that re-initializing (resetting to their initial random values) the active parameters leads to worse generalization. Further, we show that most of the active parameters are in the bottom layers, close to the input, especially as the networks become wider. Based on such observations, we study static Layer-Wise Sparse (LWS) SGD, which only updates some subsets of layers. We find that only updating the top and bottom layers have good generalization and, as expected, only updating the top layers yields a fast algorithm. Inspired by this, we investigate probabilistic LWS-SGD, which mostly updates the top layers and occasionally updates the full network. We show that probabilistic LWS-SGD matches the generalization performance of vanilla SGD and the back-propagation time can be 2-5 times more efficient.

📄 PDF Abstract BibTeX arXiv:2102.13522

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningInductive Bias

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Kernel and Rich Regimes in Overparametrized Models

2019-06-13 · Blake Woodworth, Suriya Gunasekar, Pedro Savarese, Edward Moroshko 외

A recent line of work studies overparametrized neural networks in the "kernel regime," i.e. when the network behaves during training as a kernelized linear predictor, and thus training with gradient descent has the effec…

Kernel and Rich Regimes in Overparametrized Models

2020-02-20 · Blake Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko 외

A recent line of work studies overparametrized neural networks in the "kernel regime," i.e. when the network behaves during training as a kernelized linear predictor, and thus training with gradient descent has the effec…

On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime

2026-01-06 · Shuai Jiang, Alexey Voronin, Eric Cyr, Ben Southworth arxiv

Spectral bias, the tendency of neural networks to learn low frequencies first, can be both a blessing and a curse. While it enhances the generalization capabilities by suppressing high-frequency noise, it can be a limita…

Optimal Testing in the Experiment-rich Regime

2018-05-30 · Sven Schmit, Virag Shah, Ramesh Johari

Motivated by the widespread adoption of large-scale A/B testing in industry, we propose a new experimentation framework for the setting where potential experiments are abundant (i.e., many hypotheses are available to tes…

Experimental Design

Regime-Adaptive Bayesian Optimization via Dirichlet Process Mixtures of Gaussian Processes

2026-01-27 · Yan Zhang, Xuefeng Liu, Sipeng Chen, Sascha Ranftl 외 arxiv

Standard Bayesian Optimization (BO) assumes uniform smoothness across the search space an assumption violated in multi-regime problems such as molecular conformation search through distinct energy basins or drug discover…

Gaussian ProcessesDrug Discovery