paper-with-me

홈 › Papers

Untangling tradeoffs between recurrence and self-attention in neural networks

2020-06-16 · Giancarlo Kerg, Bhargav Kanuparthi, Anirudh Goyal, Kyle Goyette, Yoshua Bengio, Guillaume Lajoie

Attention and self-attention mechanisms, are now central to state-of-the-art deep learning on sequential tasks. However, most recent progress hinges on heuristic approaches with limited understanding of attention's role in model optimization and computation, and rely on considerable memory and computational resources that scale poorly. In this work, we present a formal analysis of how self-attention affects gradient propagation in recurrent networks, and prove that it mitigates the problem of vanishing gradients when trying to capture long-term dependencies by establishing concrete bounds for gradient norms. Building on these results, we propose a relevancy screening mechanism, inspired by the cognitive process of memory consolidation, that allows for a scalable use of sparse self-attention with recurrence. While providing guarantees to avoid vanishing gradients, we use simple numerical experiments to demonstrate the tradeoffs in performance and computational resources by efficiently balancing attention and recurrence. Based on our results, we propose a concrete direction of research to improve scalability of attentive networks.

📄 PDF Abstract BibTeX arXiv:2006.09471

Code (0)

등록된 구현이 없습니다.

Tasks

Model Optimization

Similar Papers 제목 키워드 기반

Untangling tradeoffs between recurrence and self-attention in artificial neural networks

2020-12-01 · NeurIPS 2020 12 · Giancarlo Kerg, Bhargav Kanuparthi, Anirudh Goyal Alias Parth Goyal, Kyle Goyette 외

Attention and self-attention mechanisms, are now central to state-of-the-art deep learning on sequential tasks. However, most recent progress hinges on heuristic approaches with limited understanding of attention's role …

Model Optimization

Untangling Dense Non-Planar Knots by Learning Manipulation Features and Recovery Policies

2021-06-29 · Priya Sundaresan, Jennifer Grannen, Brijen Thananjeyan, Ashwin Balakrishna 외

Robot manipulation for untangling 1D deformable structures such as ropes, cables, and wires is challenging due to their infinite dimensional configuration space, complex dynamics, and tendency to self-occlude. Analytical…

Robot Manipulation

Untangling Dense Knots by Learning Task-Relevant Keypoints

2020-11-10 · Jennifer Grannen, Priya Sundaresan, Brijen Thananjeyan, Jeffrey Ichnowski 외

Untangling ropes, wires, and cables is a challenging task for robots due to the high-dimensional configuration space, visual homogeneity, self-occlusions, and complex dynamics. We consider dense (tight) knots that lack s…

NRTR: A No-Recurrence Sequence-to-Sequence Model For Scene Text Recognition

2018-06-04 · Fenfen Sheng, Zhineng Chen, Bo Xu

Scene text recognition has attracted a great many researches due to its importance to various applications. Existing methods mainly adopt recurrence or convolution based networks. Though have obtained good performance, t…

DecoderOptical Character Recognition (OCR)Scene Text Recognition

Sprucing up Supersenses: Untangling the Semantic Clusters of Accompaniment and Purpose

2020-12-01 · COLING (LAW) 2020 12 · Jena D. Hwang, Nathan Schneider, Vivek Srikumar

We reevaluate an existing adpositional annotation scheme with respect to two thorny semantic domains: accompaniment and purpose. ‘Accompaniment’ broadly speaking includes two entities situated together or participating i…