paper-with-me

홈 › Papers

RRA: Recurrent Residual Attention for Sequence Learning

2017-09-12 · Cheng Wang

In this paper, we propose a recurrent neural network (RNN) with residual attention (RRA) to learn long-range dependencies from sequential data. We propose to add residual connections across timesteps to RNN, which explicitly enhances the interaction between current state and hidden states that are several timesteps apart. This also allows training errors to be directly back-propagated through residual connections and effectively alleviates gradient vanishing problem. We further reformulate an attention mechanism over residual connections. An attention gate is defined to summarize the individual contribution from multiple previous hidden states in computing the current state. We evaluate RRA on three tasks: the adding problem, pixel-by-pixel MNIST classification and sentiment analysis on the IMDB dataset. Our experiments demonstrate that RRA yields better performance, faster convergence and more stable training compared to a standard LSTM network. Furthermore, RRA shows highly competitive performance to the state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:1709.03714

Code (1)

JRC1995/Abstractive-Summarization tf

Tasks

Sentiment Analysis

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling

2025-10-02 · Daniel Gallo Fernández arxiv

Sequence-to-sequence models have become central in Artificial Intelligence, particularly following the introduction of the transformer architecture. While initially developed for Natural Language Processing, these models…

Residual Attention Net for Superior Cross-Domain Time Sequence Modeling

2020-01-13 · Seth H. Huang, Xu Lingjie, Jiang Congwei

We present a novel architecture, residual attention net (RAN), which merges a sequence architecture, universal transformer, and a computer vision architecture, residual net, with a high-way architecture for cross-domain …

Self-Attentive Residual Decoder for Neural Machine Translation

2017-09-14 · NAACL 2018 6 · Lesly Miculicich Werlen, Nikolaos Pappas, Dhananjay Ram, Andrei Popescu-Belis

Neural sequence-to-sequence networks with attention have achieved remarkable performance for machine translation. One of the reasons for their effectiveness is their ability to capture relevant source-side contextual inf…

DecoderMachine TranslationTranslation

Temporal Convolutional Attention-based Network For Sequence Modeling

2020-02-28 · Hongyan Hao, Yan Wang, Siqiao Xue, Yudi Xia 외

With the development of feed-forward models, the default model for sequence modeling has gradually evolved to replace recurrent networks. Many powerful feed-forward models based on convolutional networks and attention me…

Co-Stack Residual Affinity Networks with Multi-level Attention Refinement for Matching Text Sequences

2018-10-06 · EMNLP 2018 10 · Yi Tay, Luu Anh Tuan, Siu Cheung Hui

Learning a matching function between two text sequences is a long standing problem in NLP research. This task enables many potential applications such as question answering and paraphrase identification. This paper propo…

Paraphrase IdentificationQuestion Answering