paper-with-me

홈 › Papers

Beyond exploding and vanishing gradients: analysing RNN training using attractors and smoothness

2019-06-20 · Antônio H. Ribeiro, Koen Tiels, Luis A. Aguirre, Thomas B. Schön

The exploding and vanishing gradient problem has been the major conceptual principle behind most architecture and training improvements in recurrent neural networks (RNNs) during the last decade. In this paper, we argue that this principle, while powerful, might need some refinement to explain recent developments. We refine the concept of exploding gradients by reformulating the problem in terms of the cost function smoothness, which gives insight into higher-order derivatives and the existence of regions with many close local minima. We also clarify the distinction between vanishing gradients and the need for the RNN to learn attractors to fully use its expressive power. Through the lens of these refinements, we shed new light on recent developments in the RNN field, namely stable RNN and unitary (or orthogonal) RNNs.

📄 PDF Abstract BibTeX arXiv:1906.08482

Code (1)

antonior92/attractors-and-smoothness-RNN 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Vanishing Nodes: Another Phenomenon That Makes Training Deep Neural Networks Difficult

2019-10-22 · Wen-Yu Chang, Tsung-Nan Lin

It is well known that the problem of vanishing/exploding gradients is a challenge when training deep networks. In this paper, we describe another phenomenon, called vanishing nodes, that also increases the difficulty of …

On the difficulty of training Recurrent Neural Networks

2012-11-21 · Razvan Pascanu, Tomas Mikolov, Yoshua Bengio

There are two widely known issues with properly training Recurrent Neural Networks, the vanishing and the exploding gradient problems detailed in Bengio et al. (1994). In this paper we attempt to improve the understandin…

Exploding and vanishing gradients in deep neural networks: the effect of residual connections

2026-06-15 · Vivek S Borkar arxiv

The well known phenomenon of exploding and vanishing gradients in deep neural networks is analyzed using multiplicative ergodic theory. The effect of adding a residual connection is explained in this context. Specificall…

A unified framework for Hamiltonian deep neural networks

2021-04-27 · Clara L. Galimberti, Liang Xu, Giancarlo Ferrari Trecate

Training deep neural networks (DNNs) can be difficult due to the occurrence of vanishing/exploding gradients during weight optimization. To avoid this problem, we propose a class of DNNs stemming from the time discretiza…

Deeply Shared Filter Bases for Parameter-Efficient Convolutional Neural Networks

2020-06-09 · NeurIPS 2021 12 · Woochul Kang, Daeyeon Kim

Modern convolutional neural networks (CNNs) have massive identical convolution blocks, and, hence, recursive sharing of parameters across these blocks has been proposed to reduce the amount of parameters. However, naive …

image-classificationImage Classificationobject-detectionObject Detection