paper-with-me

홈 › Papers

Unifying Top-down and Bottom-up for Recurrent Visual Attention

2021-09-29 · Gang Chen

The idea of using the recurrent neural network for visual attention has gained popularity in computer vision community. Although the recurrent visual attention model (RAM) leverages the glimpses with more large patch size to increasing its scope, it may result in high variance and instability. For example, we need the Gaussian policy with high variance to explore object of interests in a large image, which may cause randomized search and unstable learning. In this paper, we propose to unify the top-down and bottom-up attention together for recurrent visual attention. Our model exploits the image pyramids and Q-learning to select regions of interests in the top-down attention mechanism, which in turn to guide the policy search in the bottom-up approach. In addition, we add another two constraints over the bottom-up recurrent neural networks for better exploration. We train our model in an end-to-end reinforcement learning framework, and evaluate our method on visual classification tasks. The experimental results outperform convolutional neural networks (CNNs) baseline and the bottom-up recurrent models with visual attention.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Q-Learning

Methods 이 논문이 사용한 방법론

Q-Learning Q-Learning is an off-policy temporal difference control algorithm: $$Q\left(S\_{t}, A\_{t}\right) \leftarrow Q\left(S\_{t}, A\_{t}\right) + \alpha\left[R_{t+1} +…

Similar Papers 제목 키워드 기반

Where to Look: A Unified Attention Model for Visual Recognition with Reinforcement Learning

2021-11-13 · Gang Chen

The idea of using the recurrent neural network for visual attention has gained popularity in computer vision community. Although the recurrent attention model (RAM) leverages the glimpses with more large patch size to in…

Q-LearningReinforcement Learning (RL)

Knowledge Guided Bidirectional Attention Network for Human-Object Interaction Detection

2022-07-16 · Jingjia Huang, Baixiang Yang

Human Object Interaction (HOI) detection is a challenging task that requires to distinguish the interaction between a human-object pair. Attention based relation parsing is a popular and effective strategy utilized in HO…

DecoderHuman-Object Interaction DetectionRelation

Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers

2023-03-16 · CVPR 2024 1 · Zhibo Yang, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue 외

Most models of visual attention aim at predicting either top-down or bottom-up control, as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT), a s…

Scanpath prediction

Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over Modules

2020-06-30 · ICML 2020 1 · Sarthak Mittal, Alex Lamb, Anirudh Goyal, Vikram Voleti 외

Robust perception relies on both bottom-up and top-down signals. Bottom-up signals consist of what's directly observed through sensation. Top-down signals consist of beliefs and expectations based on past experience and …

image-classificationImage ClassificationLanguage ModelingLanguage Modelling+3

Deep Predictive Coding Network with Local Recurrent Processing for Object Recognition

2018-05-19 · NeurIPS 2018 12 · Kuan Han, Haiguang Wen, Yizhen Zhang, Di Fu 외

Inspired by "predictive coding" - a theory in neuroscience, we develop a bi-directional and dynamic neural network with local recurrent processing, namely predictive coding network (PCN). Unlike feedforward-only convolut…

image-classificationImage ClassificationObject RecognitionPrediction