paper-with-me

홈 › Papers

Simple Local Attentions Remain Competitive for Long-Context Tasks

2021-12-14 · NAACL 2022 7 · Wenhan Xiong, Barlas Oğuz, Anchit Gupta, Xilun Chen, Diana Liskovich, Omer Levy, Wen-tau Yih, Yashar Mehdad

Many NLP tasks require processing long contexts beyond the length limit of pretrained models. In order to scale these models to longer text sequences, many efficient long-range attention variants have been proposed. Despite the abundance of research along this direction, it is still difficult to gauge the relative effectiveness of these models in practical use cases, e.g., if we apply these models following the pretrain-and-finetune paradigm. In this work, we aim to conduct a thorough analysis of these emerging models with large-scale and controlled experiments. For each attention variant, we pretrain large-size models using the same long-doc corpus and then finetune these models for real-world long-context tasks. Our findings reveal pitfalls of an existing widely-used long-range benchmark and show none of the tested efficient attentions can beat a simple local window attention under standard pretraining paradigms. Further analysis on local attention variants suggests that even the commonly used attention-window overlap is not necessary to achieve good downstream results -- using disjoint local attentions, we are able to build a simpler and more efficient long-doc QA model that matches the performance of Longformer~\citep{longformer} with half of its pretraining compute. The code to replicate our experiments can be found at https://github.com/pytorch/fairseq/tree/main/examples/xformers

📄 PDF Abstract BibTeX arXiv:2112.07210

Code (1)

pytorch/fairseq 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Simple Local Attentions Remain Competitive for Long-Context Tasks

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Many NLP tasks require processing long contexts beyond the length limit of pretrained models. In order to scale these models to longer text sequences, many efficient long-range attention variants have been proposed. Desp…

Efficient Attentions for Long Document Summarization

2021-04-05 · NAACL 2021 4 · Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji 외

The quadratic computational and memory complexities of large Transformers have limited their scalability for long document summarization. In this paper, we propose Hepos, a novel efficient encoder-decoder attention with …

DecoderDocument Summarization

Learning Deep Local Features With Multiple Dynamic Attentions for Large-Scale Image Retrieval

2021-01-01 · ICCV 2021 10 · Hui Wu, Min Wang, Wengang Zhou, Houqiang Li

In image retrieval, learning local features with deep convolutional networks has been demonstrated effective to improve the performance. To discriminate deep local features, some research efforts turn to attention le…

Image RetrievalMetric LearningRetrieval

Interweaved Graph and Attention Network for 3D Human Pose Estimation

2023-04-27 · Ti Wang, Hong Liu, Runwei Ding, Wenhao Li 외

Despite substantial progress in 3D human pose estimation from a single-view image, prior works rarely explore global and local correlations, leading to insufficient learning of human skeleton representations. To address …

3D Human Pose EstimationPose Estimation

Generative Flows with Invertible Attentions

2021-06-07 · CVPR 2022 1 · Rhea Sanjay Sukthanker, Zhiwu Huang, Suryansh Kumar, Radu Timofte 외

Flow-based generative models have shown an excellent ability to explicitly learn the probability density function of data via a sequence of invertible transformations. Yet, learning attentions in generative flows remains…

Image Generation