paper-with-me

홈 › Papers

On the difficulty of training Recurrent Neural Networks

2012-11-21 · Razvan Pascanu, Tomas Mikolov, Yoshua Bengio

There are two widely known issues with properly training Recurrent Neural Networks, the vanishing and the exploding gradient problems detailed in Bengio et al. (1994). In this paper we attempt to improve the understanding of the underlying issues by exploring these problems from an analytical, a geometric and a dynamical systems perspective. Our analysis is used to justify a simple yet effective solution. We propose a gradient norm clipping strategy to deal with exploding gradients and a soft constraint for the vanishing gradients problem. We validate empirically our hypothesis and proposed solutions in the experimental section.

📄 PDF Abstract BibTeX arXiv:1211.5063

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive recurrent vision performs zero-shot computation scaling to unseen difficulty levels

2023-11-12 · NeurIPS 2023 11

Humans solving algorithmic (or) reasoning problems typically exhibit solution times that grow as a function of problem difficulty. Adaptive recurrent neural networks have been shown to exhibit this property for various l…

PathfinderVisual ReasoningZero-shot Generalization

Recurrent Affine Transformation for Text-to-image Synthesis

2022-04-22 · Senmao Ye, Fei Liu, Minkui Tan

Text-to-image synthesis aims to generate natural images conditioned on text descriptions. The main difficulty of this task lies in effectively fusing text information into the image synthesis process. Existing methods us…

Text-to-Image Generation

Simple Recurrent Units for Highly Parallelizable Recurrence

2017-09-08 · EMNLP 2018 10 · Tao Lei, Yu Zhang, Sida I. Wang, Hui Dai 외

Common recurrent neural architectures scale poorly due to the intrinsic difficulty in parallelizing their state computations. In this work, we propose the Simple Recurrent Unit (SRU), a light recurrent unit that balances…

General ClassificationMachine TranslationQuestion AnsweringText Classification+1

Training RNNs as Fast as CNNs

2018-01-01 · ICLR 2018 1 · Tao Lei, Yu Zhang, Yoav Artzi

Common recurrent neural network architectures scale poorly due to the intrinsic difficulty in parallelizing their state computations. In this work, we propose the Simple Recurrent Unit (SRU) architecture, a recurrent uni…

General ClassificationLanguage ModelingLanguage ModellingQuestion Answering+3

Recurrent Reasoning on Symbolic Puzzles with Sequence Models

2026-04-19 · Gowrav Mannem, Chowdhury Marzia Mahjabin, Jason Chen, Shivank Garg 외 arxiv

Large language models often appear strong on symbolic and algorithmic tasks, yet this apparent strength can hide brittle behaviour when problems become longer, harder, or slightly out of distribution. A major limitation …