SALYPATH: A Deep-Based Architecture for visual attention prediction
Human vision is naturally more attracted by some regions within their field of view than others. This intrinsic selectivity mechanism, so-called visual attention, is influenced by both high- and low-level factors; such as the global environment (illumination, background texture, etc.), stimulus characteristics (color, intensity, orientation, etc.), and some prior visual information. Visual attention is useful for many computer vision applications such as image compression, recognition, and captioning. In this paper, we propose an end-to-end deep-based method, so-called SALYPATH (SALiencY and scanPATH), that efficiently predicts the scanpath of an image through features of a saliency model. The idea is predict the scanpath by exploiting the capacity of a deep-based model to predict the saliency. The proposed method was evaluated through 2 well-known datasets. The results obtained showed the relevance of the proposed framework comparing to state-of-the-art models.
Code (0)
등록된 구현이 없습니다.
Tasks
Saliency PredictionScanpath predictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SalyPath360: Saliency and Scanpath Prediction Framework for Omnidirectional Images
This paper introduces a new framework to predict visual attention of omnidirectional images. The key setup of our architecture is the simultaneous prediction of the saliency map and a corresponding scanpath for a given s…
DecoderPredictionScanpath predictionSaliency-based Sequential Image Attention with Multiset Prediction
Humans process visual scenes selectively and sequentially using attention. Central to models of human visual attention is the saliency map. We propose a hierarchical visual architecture that operates on a saliency map an…
ClassificationGeneral Classificationimage-classificationImage Classification+5Improved Fusion of Visual and Language Representations by Dense Symmetric Co-Attention for Visual Question Answering
A key solution to visual question answering (VQA) exists in how to fuse visual and language features extracted from an input image and question. We show that an attention mechanism that enables dense, bi-directional inte…
Visual Question AnsweringVisual Question Answering (VQA)Deep learning investigation for chess player attention prediction using eye-tracking and game data
This article reports on an investigation of the use of convolutional neural networks to predict the visual attention of chess players. The visual attention model described in this article has been created to generate sal…
DecoderContext-empowered Visual Attention Prediction in Pedestrian Scenarios
Effective and flexible allocation of visual attention is key for pedestrians who have to navigate to a desired goal under different conditions of urgency and safety preferences. While automatic modelling of pedestrian at…
DecoderNavigatePredictionSaliency Prediction