paper-with-me

Papers

Contextual Encoder-Decoder Network for Visual Saliency Prediction

2019-02-18 · Alexander Kroner, Mario Senden, Kurt Driessens, Rainer Goebel

Predicting salient regions in natural images requires the detection of objects that are present in a scene. To develop robust representations for this challenging task, high-level visual features at multiple spatial scales must be extracted and augmented with contextual information. However, existing models aimed at explaining human fixation maps do not incorporate such a mechanism explicitly. Here we propose an approach based on a convolutional neural network pre-trained on a large-scale image classification task. The architecture forms an encoder-decoder structure and includes a module with multiple convolutional layers at different dilation rates to capture multi-scale features in parallel. Moreover, we combine the resulting representations with global scene information for accurately predicting visual saliency. Our model achieves competitive and consistent results across multiple evaluation metrics on two public saliency benchmarks and we demonstrate the effectiveness of the suggested approach on five datasets and selected examples. Compared to state of the art approaches, the network is based on a lightweight image classification backbone and hence presents a suitable choice for applications with limited computational resources, such as (virtual) robotic systems, to estimate human fixations across complex natural scenes.

📄 PDF Abstract BibTeX arXiv:1902.06634

Code (4)

alexanderkroner/saliency 공식 구현 tf
gradio-app/saliency tf
nvinden/7ChannelEML tf
yahsieh37/Visual-Saliency-Prediction tf

Tasks

DecoderGeneral Classificationimage-classificationImage ClassificationPredictionSaliency PredictionVideo Saliency Detection

Similar Papers 제목 키워드 기반

DAVE: A Deep Audio-Visual Embedding for Dynamic Saliency Prediction

2019-05-25 · Hamed R. -Tavakoli, Ali Borji, Esa Rahtu, Juho Kannala

This paper studies audio-visual deep saliency prediction. It introduces a conceptually simple and effective Deep Audio-Visual Embedding for dynamic saliency prediction dubbed ``DAVE" in conjunction with our efforts towar…

DecoderPredictionSaliency PredictionVideo Saliency Prediction

ViNet: Pushing the limits of Visual Modality for Audio-Visual Saliency Prediction

2020-12-11 · Samyak Jain, Pradeep Yarlagadda, Shreyank Jyoti, Shyamgopal Karthik 외

We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network trained for action recognition, and the…

Action RecognitionDecoderPredictionSaliency Prediction+2

SalyPath360: Saliency and Scanpath Prediction Framework for Omnidirectional Images

2022-01-01 · Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Mohamed Sayeh

This paper introduces a new framework to predict visual attention of omnidirectional images. The key setup of our architecture is the simultaneous prediction of the saliency map and a corresponding scanpath for a given s…

DecoderPredictionScanpath prediction

MDS-ViTNet: Improving saliency prediction for Eye-Tracking with Vision Transformer

2024-05-29 · Polezhaev Ignat, Goncharenko Igor, Iurina Natalya

In this paper, we present a novel methodology we call MDS-ViTNet (Multi Decoder Saliency by Vision Transformer Network) for enhancing visual saliency prediction or eye-tracking. This approach holds significant potential …

DecoderMarketingSaliency PredictionTransfer Learning

Context-empowered Visual Attention Prediction in Pedestrian Scenarios

2022-10-30 · Igor Vozniak, Philipp Mueller, Lorena Hell, Nils Lipp 외

Effective and flexible allocation of visual attention is key for pedestrians who have to navigate to a desired goal under different conditions of urgency and safety preferences. While automatic modelling of pedestrian at…

DecoderNavigatePredictionSaliency Prediction