paper-with-me

홈 › Papers

Rethinking 360deg Image Visual Attention Modelling With Unsupervised Learning.

2021-01-01 · ICCV 2021 10 · Yasser Abdelaziz Dahou Djilali, Tarun Krishna, Kevin McGuinness, Noel E. O'Connor

Despite the success of self-supervised representation learning on planar data, to date it has not been studied on 360deg images. In this paper, we extend recent advances in contrastive learning to learn latent representations that are sufficiently invariant to be highly effective for spherical saliency prediction as a downstream task. We argue that omni-directional images are particularly suited to such an approach due to the geometry of the data domain. To verify this hypothesis, we design an unsupervised framework that effectively maximizes the mutual information between the different views from both the equator and the poles. We show that the decoder is able to learn good quality saliency distributions from the encoder embeddings. Our model compares favorably with fully-supervised learning methods on the Salient360!, VR-EyeTracking and Sitzman datasets. This performance is achieved using an encoder that is trained in a completely unsupervised way and a relatively lightweight supervised decoder (3.8 X fewer parameters in the case of the ResNet50 encoder). We believe that this combination of supervised and unsupervised learning is an important step toward flexible formulations of human visual attention.

📄 PDF Abstract BibTeX

Code (1)

KrishnaTarun/360_unsupervised_saliency 공식 구현 pytorch

Tasks

Contrastive LearningDecoderRepresentation LearningSaliency Prediction

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Rethinking Alignment and Uniformity in Unsupervised Semantic Segmentation

2022-11-26 · Daoan Zhang, Chenming Li, Haoquan Li, Wenjian Huang 외

Unsupervised image semantic segmentation(UISS) aims to match low-level visual features with semantic-level representations without outer supervision. In this paper, we address the critical properties from the view of fea…

Representation LearningSegmentationSemantic SegmentationUnsupervised Image Segmentation+1

Self Supervised Scanpath Prediction Framework for Painting Images

2022-06-19 · CVPR 2022 6 · Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, Alessandro Bruno

In our paper, we propose a novel strategy to learn distortion invariant latent representation from painting pictures for visual attention modelling downstream task. In further detail, we design an unsupervised framework …

DecoderPredictionScanpath prediction

Masked Feature Modelling: Feature Masking for the Unsupervised Pre-training of a Graph Attention Network Block for Bottom-up Video Event Recognition

2023-08-24 · Dimitrios Daskalakis, Nikolaos Gkalelis, Vasileios Mezaris

In this paper, we introduce Masked Feature Modelling (MFM), a novel approach for the unsupervised pre-training of a Graph Attention Network (GAT) block. MFM utilizes a pretrained Visual Tokenizer to reconstruct masked fe…

Graph AttentionUnsupervised Pre-training

Looking to Learn: Token-wise Dynamic Gating for Low-Resource Vision-Language Modelling

2025-10-09 · Bianca-Mihaela Ganescu, Suchir Salhan, Andrew Caines, Paula Buttery arxiv

Training vision-language models on cognitively-plausible amounts of data requires rethinking how models integrate multimodal information. Within the constraints of the Vision track for the BabyLM Challenge 2025, we propo…

Language ModellingVisual Grounding

Trends, Applications, and Challenges in Human Attention Modelling

2024-02-28 · Giuseppe Cartella, Marcella Cornia, Vittorio Cuculo, Alessandro D'Amelio 외

Human attention modelling has proven, in recent years, to be particularly useful not only for understanding the cognitive processes underlying visual exploration, but also for providing support to artificial intelligence…

Language Modelling