paper-with-me

홈 › Papers

Training Multi-Layer Over-Parametrized Neural Network in Subquadratic Time

2021-12-14 · Zhao Song, Lichen Zhang, Ruizhe Zhang

We consider the problem of training a multi-layer over-parametrized neural network to minimize the empirical risk induced by a loss function. In the typical setting of over-parametrization, the network width $m$ is much larger than the data dimension $d$ and the number of training samples $n$ ($m=\mathrm{poly}(n,d)$), which induces a prohibitive large weight matrix $W\in \mathbb{R}^{m\times m}$ per layer. Naively, one has to pay $O(m^2)$ time to read the weight matrix and evaluate the neural network function in both forward and backward computation. In this work, we show how to reduce the training cost per iteration. Specifically, we propose a framework that uses $m^2$ cost only in the initialization phase and achieves \emph{a truly subquadratic cost per iteration} in terms of $m$, i.e., $m^{2-\Omega(1)}$ per iteration. Our result has implications beyond standard over-parametrization theory, as it can be viewed as designing an efficient data structure on top of a pre-trained large model to further speed up the fine-tuning process, a core procedure to deploy large language models (LLM).

📄 PDF Abstract BibTeX arXiv:2112.07628

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Linear Convergence of SGD on Overparametrized Shallow Neural Networks

2021-09-29 · Paul Rolland, Ali Ramezani-Kebrya, ChaeHwan Song, Fabian Latorre 외

Despite the non-convex landscape, first-order methods can be shown to reach global minima when training overparameterized neural networks, where the number of parameters far exceed the number of training data. In this wo…

Hyena Hierarchy: Towards Larger Convolutional Language Models

2023-02-21 · Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu 외

Recent advances in deep learning have relied heavily on the use of large Transformers due to their ability to learn at scale. However, the core building block of Transformers, the attention operator, exhibits quadratic c…

2k8kLanguage ModelingLanguage Modelling+1

Beyond the Quadratic Approximation: the Multiscale Structure of Neural Network Loss Landscapes

2022-04-24 · Chao Ma, Daniel Kunin, Lei Wu, Lexing Ying

A quadratic approximation of neural network loss landscapes has been extensively used to study the optimization process of these networks. Though, it usually holds in a very small neighborhood of the minimum, it cannot e…

Subquadratic Overparameterization for Shallow Neural Networks

2021-11-02 · NeurIPS 2021 12 · ChaeHwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari 외

Overparameterization refers to the important phenomenon where the width of a neural network is chosen such that learning algorithms can provably attain zero loss in nonconvex training. The existing theory establishes suc…

Kernel and Rich Regimes in Overparametrized Models

2019-06-13 · Blake Woodworth, Suriya Gunasekar, Pedro Savarese, Edward Moroshko 외

A recent line of work studies overparametrized neural networks in the "kernel regime," i.e. when the network behaves during training as a kernelized linear predictor, and thus training with gradient descent has the effec…