paper-with-me

홈 › Papers

Feed-Forward Networks with Attention Can Solve Some Long-Term Memory Problems

2015-12-29 · Colin Raffel, Daniel P. W. Ellis

We propose a simplified model of attention which is applicable to feed-forward neural networks and demonstrate that the resulting model can solve the synthetic "addition" and "multiplication" long-term memory problems for sequence lengths which are both longer and more widely varying than the best published results for these tasks.

📄 PDF Abstract BibTeX arXiv:1512.08756

Code (5)

WenYanger/Contextual-Attention pytorch
dtsbourg/ff-attention pytorch
firfre/Chinese-Text-sentiment-analysis-with-key-word tf
shawnyxiao/textclassification-keras tf
zaczou/keras_summary tf

Similar Papers 제목 키워드 기반

Feedback Attention for Cell Image Segmentation

2020-08-14 · Hiroki Tsuda, Eisuke Shibuya, Kazuhiro Hotta

In this paper, we address cell image segmentation task by Feedback Attention mechanism like feedback processing. Unlike conventional neural network models of feedforward processing, we focused on the feedback processing …

Image SegmentationSegmentationSemantic Segmentation

Integrating Quantum-Classical Attention in Patch Transformers for Enhanced Time Series Forecasting

2025-03-31 · Sanjay Chakraborty, Fredrik Heintz

QCAAPatchTF is a quantum attention network integrated with an advanced patch-based transformer, designed for multivariate time series forecasting, classification, and anomaly detection. Leveraging quantum superpositions,…

Anomaly DetectionMultivariate Time Series ForecastingTime SeriesTime Series Forecasting

Augmenting Self-attention with Persistent Memory

2019-07-02 · Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou 외

Transformer networks have lead to important progress in language modeling and machine translation. These models include two consecutive modules, a feed-forward layer and a self-attention layer. The latter allows the netw…

Language ModelingLanguage ModellingTranslation

An approach to reachability analysis for feed-forward ReLU neural networks

2017-06-22 · Alessio Lomuscio, Lalit Maganti

We study the reachability problem for systems implemented as feed-forward neural networks whose activation function is implemented via ReLU functions. We draw a correspondence between establishing whether some arbitrary …

CoLT5: Faster Long-Range Transformers with Conditional Computation

2023-03-17 · Joshua Ainslie, Tao Lei, Michiel de Jong, Santiago Ontañón 외

Many natural language processing tasks benefit from long inputs, but processing long documents with Transformers is expensive -- not only due to quadratic attention complexity but also from applying feedforward and proje…

Long-range modeling