paper-with-me

홈 › Papers

On the Convergence Rate of Training Recurrent Neural Networks

2018-10-29 · NeurIPS 2019 12 · Zeyuan Allen-Zhu, Yuanzhi Li, Zhao Song

How can local-search methods such as stochastic gradient descent (SGD) avoid bad local minima in training multi-layer neural networks? Why can they fit random labels even given non-convex and non-smooth architectures? Most existing theory only covers networks with one hidden layer, so can we go deeper? In this paper, we focus on recurrent neural networks (RNNs) which are multi-layer networks widely used in natural language processing. They are harder to analyze than feedforward neural networks, because the $\textit{same}$ recurrent unit is repeatedly applied across the entire time horizon of length $L$, which is analogous to feedforward networks of depth $L$. We show when the number of neurons is sufficiently large, meaning polynomial in the training data size and in $L$, then SGD is capable of minimizing the regression loss in the linear convergence rate. This gives theoretical evidence of how RNNs can memorize data. More importantly, in this paper we build general toolkits to analyze multi-layer networks with ReLU activations. For instance, we prove why ReLU activations can prevent exponential gradient explosion or vanishing, and build a perturbation theory to analyze first-order approximation of multi-layer networks.

📄 PDF Abstract BibTeX arXiv:1810.12065

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Deep Recurrent Convolutional Neural Network: Improving Performance For Speech Recognition

2016-11-22 · Zewang Zhang, Zheng Sun, Jiaqi Liu, Jingwen Chen 외

A deep learning approach has been widely applied in sequence modeling problems. In terms of automatic speech recognition (ASR), its performance has significantly been improved by increasing large speech corpus and deeper…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Coordinate descent on the orthogonal group for recurrent neural network training

2021-07-30 · Estelle Massart, Vinayak Abrol

We propose to use stochastic Riemannian coordinate descent on the orthogonal group for recurrent neural network training. The algorithm rotates successively two columns of the recurrent matrix, an operation that can be e…

Evaluating the Stability of Recurrent Neural Models during Training with Eigenvalue Spectra Analysis

2019-05-08 · Priyadarshini Panda, Efstathia Soufleri, Kaushik Roy

We analyze the stability of recurrent networks, specifically, reservoir computing models during training by evaluating the eigenvalue spectra of the reservoir dynamics. To circumvent the instability arising in examining …

regressionvalid

Block-Recurrent Dynamics in Vision Transformers

2025-12-23 · Mozes Jacobs, Thomas Fel, Richard Hakim, Alessandra Brondetta 외 arxiv

As Vision Transformers (ViTs) become standard vision backbones, a mechanistic account of their computational phenomenology is essential. Despite architectural cues that hint at dynamical structure, there is no settled fr…

Convergence Analysis of Real-time Recurrent Learning (RTRL) for a class of Recurrent Neural Networks

2025-01-14 · Samuel Chun-Hei Lam, Justin Sirignano, Konstantinos Spiliopoulos

Recurrent neural networks (RNNs) are commonly trained with the truncated backpropagation-through-time (TBPTT) algorithm. For the purposes of computational tractability, the TBPTT algorithm truncates the chain rule and ca…