paper-with-me

Papers

Self-Supervised Visual Representations for Cross-Modal Retrieval

2019-01-31 · Yash Patel, Lluis Gomez, Marçal Rusiñol, Dimosthenis Karatzas, C. V. Jawahar

Cross-modal retrieval methods have been significantly improved in last years with the use of deep neural networks and large-scale annotated datasets such as ImageNet and Places. However, collecting and annotating such datasets requires a tremendous amount of human effort and, besides, their annotations are usually limited to discrete sets of popular visual classes that may not be representative of the richer semantics found on large-scale cross-modal retrieval datasets. In this paper, we present a self-supervised cross-modal retrieval framework that leverages as training data the correlations between images and text on the entire set of Wikipedia articles. Our method consists in training a CNN to predict: (1) the semantic context of the article in which an image is more probable to appear as an illustration (global context), and (2) the semantic context of its caption (local context). Our experiments demonstrate that the proposed method is not only capable of learning discriminative visual representations for solving vision tasks like image classification and object detection, but that the learned representations are better for cross-modal retrieval when compared to supervised pre-training of the network on the ImageNet dataset.

📄 PDF Abstract BibTeX arXiv:1902.00378

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesCross-Modal Retrievalimage-classificationImage Classificationobject-detectionObject DetectionRetrieval

Similar Papers 제목 키워드 기반

Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision

2020-07-08 · Abhinav Shukla, Stavros Petridis, Maja Pantic

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and a…

Acoustic Scene ClassificationAction RecognitionScene ClassificationSelf-Supervised Learning+1

Does Visual Self-Supervision Improve Learning of Speech Representations for Emotion Recognition?

2020-05-04 · Abhinav Shukla, Stavros Petridis, Maja Pantic

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between au…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionFace Reconstruction+5

Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity

2021-11-09 · Pritam Sarkar, Ali Etemad

We present CrissCross, a self-supervised framework for learning audio-visual representations. A novel notion is introduced in our framework whereby in addition to learning the intra-modal and standard 'synchronous' cross…

Audio ClassificationRetrievalSelf-Supervised Action RecognitionSelf-Supervised Audio Classification+4

Visually Guided Self Supervised Learning of Speech Representations

2020-01-13 · Abhinav Shukla, Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis 외

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there …

Emotion RecognitionRepresentation LearningSelf-Supervised LearningSpeech Emotion Recognition+2

Investigating self-supervised representations for audio-visual deepfake detection

2025-11-21 · Dragos-Alexandru Boldisor, Stefan Smeu, Dan Oneata, Elisabeta Oneata arxiv

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried with…

DeepFake Detection