paper-with-me

홈 › Papers

Long Expressive Memory for Sequence Modeling

2021-10-10 · ICLR 2022 4 · T. Konstantin Rusch, Siddhartha Mishra, N. Benjamin Erichson, Michael W. Mahoney

We propose a novel method called Long Expressive Memory (LEM) for learning long-term sequential dependencies. LEM is gradient-based, it can efficiently process sequential tasks with very long-term dependencies, and it is sufficiently expressive to be able to learn complicated input-output maps. To derive LEM, we consider a system of multiscale ordinary differential equations, as well as a suitable time-discretization of this system. For LEM, we derive rigorous bounds to show the mitigation of the exploding and vanishing gradients problem, a well-known challenge for gradient-based recurrent sequential learning methods. We also prove that LEM can approximate a large class of dynamical systems to high accuracy. Our empirical results, ranging from image and time-series classification through dynamical systems prediction to speech recognition and language modeling, demonstrate that LEM outperforms state-of-the-art recurrent neural networks, gated recurrent units, and long short-term memory models.

📄 PDF Abstract BibTeX arXiv:2110.04744

Code (1)

tk-rusch/lem 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingSequential Image Classificationspeech-recognitionSpeech RecognitionTime SeriesTime Series AnalysisTime Series Classification

Similar Papers 제목 키워드 기반

Diffuser: Efficient Transformers with Multi-hop Attention Diffusion for Long Sequences

2022-10-21 · Aosong Feng, Irene Li, Yuang Jiang, Rex Ying

Efficient Transformers have been developed for long sequence modeling, due to their subquadratic memory and time complexity. Sparse Transformer is a popular approach to improving the efficiency of Transformers by restric…

Language ModelingLanguage Modellingtext-classificationText Classification

Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling

2024-02-01 · Mingze Wang, Weinan E

We conduct a systematic study of the approximation properties of Transformer for sequence modeling with long, sparse and complicated memory. We investigate the mechanisms through which different components of Transformer…

ATLAS: Learning to Optimally Memorize the Context at Test Time

2025-05-29 · Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri 외

Transformers have been established as the most popular backbones in sequence modeling, mainly due to their effectiveness in in-context retrieval tasks and the ability to learn at scale. Their quadratic memory and time co…

Common Sense ReasoningLanguage ModelingLanguage ModellingLong-Context Understanding

Nested Learning: The Illusion of Deep Learning Architectures

2025-12-31 · Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab Mirrokni arxiv

Despite the recent progresses, particularly in developing Language Models, there are fundamental challenges and unanswered questions about how such models can continually learn/memorize, self-improve, and find effective …

Continual Learning

Learning to (Learn at Test Time): RNNs with Expressive Hidden States

2024-07-05 · Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu 외

Self-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expressive power of their hidden state. We pr…

16k8kMambaSelf-Supervised Learning