paper-with-me

홈 › Papers

Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over Modules

2020-06-30 · ICML 2020 1 · Sarthak Mittal, Alex Lamb, Anirudh Goyal, Vikram Voleti, Murray Shanahan, Guillaume Lajoie, Michael Mozer, Yoshua Bengio

Robust perception relies on both bottom-up and top-down signals. Bottom-up signals consist of what's directly observed through sensation. Top-down signals consist of beliefs and expectations based on past experience and short-term memory, such as how the phrase `peanut butter and~...' will be completed. The optimal combination of bottom-up and top-down information remains an open question, but the manner of combination must be dynamic and both context and task dependent. To effectively utilize the wealth of potential top-down information available, and to prevent the cacophony of intermixed signals in a bidirectional architecture, mechanisms are needed to restrict information flow. We explore deep recurrent neural net architectures in which bottom-up and top-down signals are dynamically combined using attention. Modularity of the architecture further restricts the sharing and communication of information. Together, attention and modularity direct information flow, which leads to reliable performance improvements in perceptual and language tasks, and in particular improves robustness to distractions and noisy data. We demonstrate on a variety of benchmarks in language modeling, sequential image classification, video prediction and reinforcement learning that the \emph{bidirectional} information flow can improve results over strong baselines.

📄 PDF Abstract BibTeX arXiv:2006.16981

Code (1)

sarthmit/BRIMs 공식 구현 pytorch

Tasks

image-classificationImage ClassificationLanguage ModelingLanguage ModellingOpen-Ended Question AnsweringSequential Image ClassificationVideo Prediction

Similar Papers 제목 키워드 기반

Unifying Top-down and Bottom-up for Recurrent Visual Attention

2021-09-29 · Gang Chen

The idea of using the recurrent neural network for visual attention has gained popularity in computer vision community. Although the recurrent visual attention model (RAM) leverages the glimpses with more large patch siz…

Q-Learning

Where to Look: A Unified Attention Model for Visual Recognition with Reinforcement Learning

2021-11-13 · Gang Chen

The idea of using the recurrent neural network for visual attention has gained popularity in computer vision community. Although the recurrent attention model (RAM) leverages the glimpses with more large patch size to in…

Q-LearningReinforcement Learning (RL)

Image Captioning with Semantic Attention

2016-03-12 · CVPR 2016 6 · Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang 외

Automatically generating a natural language description of an image has attracted interests recently both because of its importance in practical applications and because it connects two major artificial intelligence fiel…

Image Captioning

A Novel Method to Study Bottom-up Visual Saliency and its Neural Mechanism

2016-04-13 · Cheng Chen, Xilin Zhang, Yizhou Wang, Fang Fang

In this study, we propose a novel method to measure bottom-up saliency maps of natural images. In order to eliminate the influence of top-down signals, backward masking is used to make stimuli (natural images) subjective…

Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering

2017-07-25 · CVPR 2018 6 · Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney 외

Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reason…

Image CaptioningVisual Question AnsweringVisual Question Answering (VQA)