paper-with-me

홈 › Papers

Not All Attention Is Needed: Gated Attention Network for Sequence Data

2019-12-01 · Lanqing Xue, Xiaopeng Li, Nevin L. Zhang

Although deep neural networks generally have fixed network structures, the concept of dynamic mechanism has drawn more and more attention in recent years. Attention mechanisms compute input-dependent dynamic attention weights for aggregating a sequence of hidden states. Dynamic network configuration in convolutional neural networks (CNNs) selectively activates only part of the network at a time for different inputs. In this paper, we combine the two dynamic mechanisms for text classification tasks. Traditional attention mechanisms attend to the whole sequence of hidden states for an input sentence, while in most cases not all attention is needed especially for long sequences. We propose a novel method called Gated Attention Network (GA-Net) to dynamically select a subset of elements to attend to using an auxiliary network, and compute attention weights to aggregate the selected elements. It avoids a significant amount of unnecessary computation on unattended elements, and allows the model to pay attention to important parts of the sequence. Experiments in various datasets show that the proposed method achieves better performance compared with all baseline models with global or local attention while requiring less computation and achieving better interpretability. It is also promising to extend the idea to more complex attention-based models, such as transformers and seq-to-seq models.

📄 PDF Abstract BibTeX arXiv:1912.00349

Code (1)

twisha96/Natural-Language-Processing pytorch

Tasks

AllSentencetext-classificationText Classification

Similar Papers 제목 키워드 기반

Sparse Modular Activation for Efficient Sequence Modeling

2023-06-19 · NeurIPS 2023 11 · Liliang Ren, Yang Liu, Shuohang Wang, Yichong Xu 외

Recent hybrid models combining Linear State Space Models (SSMs) with self-attention mechanisms have demonstrated impressive results across a range of sequence modeling tasks. However, current approaches apply attention m…

ChunkingLanguage ModelingLanguage ModellingLong-range modeling+1

Temporal Attention-Gated Model for Robust Sequence Classification

2016-12-01 · CVPR 2017 7 · Wenjie Pei, Tadas Baltrušaitis, David M. J. Tax, Louis-Philippe Morency

Typical techniques for sequence classification are designed for well-segmented sequences which have been edited to remove noisy or irrelevant parts. Therefore, such methods cannot be easily applied on noisy sequences exp…

ClassificationGeneral ClassificationmodelSentiment Analysis

Posterior Attention Models for Sequence to Sequence Learning

2019-05-01 · ICLR 2019 5 · Shiv Shankar, Sunita Sarawagi

Modern neural architectures critically rely on attention for mapping structured inputs to sequences. In this paper we show that prevalent attention architectures do not adequately model the dependence among the attention…

Morphological InflectionPositionTranslation

Exploring Sequence-to-Sequence Learning in Aspect Term Extraction

2019-07-01 · ACL 2019 7 · Dehong Ma, Sujian Li, Fangzhao Wu, Xing Xie 외

Aspect term extraction (ATE) aims at identifying all aspect terms in a sentence and is usually modeled as a sequence labeling problem. However, sequence labeling based methods cannot make full use of the overall meaning …

DecoderPositionSentenceTerm Extraction

Mega: Moving Average Equipped Gated Attention

2022-09-21 · Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He 외

The design choices in the Transformer attention mechanism, including weak inductive bias and quadratic computational complexity, have limited its application for modeling long sequences. In this paper, we introduce Mega,…

Image ClassificationInductive BiasLanguage ModelingLanguage Modelling+5