paper-with-me

홈 › Papers

Exploring the Long-Term Generalization of Counting Behavior in RNNs

2022-11-29 · Nadine El-Naggar, Pranava Madhyastha, Tillman Weyde

In this study, we investigate the generalization of LSTM, ReLU and GRU models on counting tasks over long sequences. Previous theoretical work has established that RNNs with ReLU activation and LSTMs have the capacity for counting with suitable configuration, while GRUs have limitations that prevent correct counting over longer sequences. Despite this and some positive empirical results for LSTMs on Dyck-1 languages, our experimental results show that LSTMs fail to learn correct counting behavior for sequences that are significantly longer than in the training data. ReLUs show much larger variance in behavior and in most cases worse generalization. The long sequence generalization is empirically related to validation loss, but reliable long sequence generalization seems not practically achievable through backpropagation with current techniques. We demonstrate different failure modes for LSTMs, GRUs and ReLUs. In particular, we observe that the saturation of activation functions in LSTMs and the correct weight setting for ReLUs to generalize counting behavior are not achieved in standard training regimens. In summary, learning generalizable counting behavior is still an open problem and we discuss potential approaches for further research.

📄 PDF Abstract BibTeX arXiv:2211.16429

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

fail 설명 없음
Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
GRU A Gated Recurrent Unit, or GRU, is a type of recurrent neural network. It is similar to an LSTM, but only has two gates - a reset…
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Discounting and Drug Seeking in Biological Hierarchical Reinforcement Learning

2025-06-05 · Vardhan Palod, Pranav Mahajan, Veeky Baths, Boris S. Gutkin

Despite a strong desire to quit, individuals with long-term substance use disorder (SUD) often struggle to resist drug use, even when aware of its harmful consequences. This disconnect between knowledge and compulsive be…

Hierarchical Reinforcement Learningreinforcement-learningReinforcement Learning

\$1 Today or \$2 Tomorrow? The Answer is in Your Facebook Likes

2017-03-22 · Tao Ding, Warren K. Bickel, SHimei Pan

In economics and psychology, delay discounting is often used to characterize how individuals choose between a smaller immediate reward and a larger delayed reward. People with higher delay discounting rate (DDR) often ch…

Decision Making

Exploring a strongly non-Markovian animal behavior

2020-12-31 · Vasyl Alba, Gordon J. Berman, William Bialek, Joshua W. Shaevitz

A freely walking fly visits roughly 100 stereotyped states in a strongly non-Markovian sequence. To explore these dynamics, we develop a generalization of the information bottleneck method, compressing the large number o…

Social Discounting and the Long Rate of Interest

2015-09-28

The well-known theorem of Dybvig, Ingersoll and Ross shows that the long zero-coupon rate can never fall. This result, which, although undoubtedly correct, has been regarded by many as surprising, stems from the implicit…

Management

Market and Long Term Accounting Operational Performance

2019-07-26

Following the value relevance literature, this study verifies whether the marketplace differentiates companies of high, medium, and low long-term operational performance, measured by accounting information on profitabili…