Exploring the Long-Term Generalization of Counting Behavior in RNNs
In this study, we investigate the generalization of LSTM, ReLU and GRU models on counting tasks over long sequences. Previous theoretical work has established that RNNs with ReLU activation and LSTMs have the capacity for counting with suitable configuration, while GRUs have limitations that prevent correct counting over longer sequences. Despite this and some positive empirical results for LSTMs on Dyck-1 languages, our experimental results show that LSTMs fail to learn correct counting behavior for sequences that are significantly longer than in the training data. ReLUs show much larger variance in behavior and in most cases worse generalization. The long sequence generalization is empirically related to validation loss, but reliable long sequence generalization seems not practically achievable through backpropagation with current techniques. We demonstrate different failure modes for LSTMs, GRUs and ReLUs. In particular, we observe that the saturation of activation functions in LSTMs and the correct weight setting for ReLUs to generalize counting behavior are not achieved in standard training regimens. In summary, learning generalizable counting behavior is still an open problem and we discuss potential approaches for further research.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Discounting and Drug Seeking in Biological Hierarchical Reinforcement Learning
Despite a strong desire to quit, individuals with long-term substance use disorder (SUD) often struggle to resist drug use, even when aware of its harmful consequences. This disconnect between knowledge and compulsive be…
Hierarchical Reinforcement Learningreinforcement-learningReinforcement Learning\$1 Today or \$2 Tomorrow? The Answer is in Your Facebook Likes
In economics and psychology, delay discounting is often used to characterize how individuals choose between a smaller immediate reward and a larger delayed reward. People with higher delay discounting rate (DDR) often ch…
Decision MakingExploring a strongly non-Markovian animal behavior
A freely walking fly visits roughly 100 stereotyped states in a strongly non-Markovian sequence. To explore these dynamics, we develop a generalization of the information bottleneck method, compressing the large number o…
Social Discounting and the Long Rate of Interest
The well-known theorem of Dybvig, Ingersoll and Ross shows that the long zero-coupon rate can never fall. This result, which, although undoubtedly correct, has been regarded by many as surprising, stems from the implicit…
ManagementMarket and Long Term Accounting Operational Performance
Following the value relevance literature, this study verifies whether the marketplace differentiates companies of high, medium, and low long-term operational performance, measured by accounting information on profitabili…