paper-with-me

홈 › Papers

Self-Supervised Learning from Web Data for Multimodal Retrieval

2019-01-07 · Raul Gomez, Lluis Gomez, Jaume Gibert, Dimosthenis Karatzas

Self-Supervised learning from multimodal image and text data allows deep neural networks to learn powerful features with no need of human annotated data. Web and Social Media platforms provide a virtually unlimited amount of this multimodal data. In this work we propose to exploit this free available data to learn a multimodal image and text embedding, aiming to leverage the semantic knowledge learnt in the text domain and transfer it to a visual model for semantic image retrieval. We demonstrate that the proposed pipeline can learn from images with associated textwithout supervision and analyze the semantic structure of the learnt joint image and text embedding space. We perform a thorough analysis and performance comparison of five different state of the art text embeddings in three different benchmarks. We show that the embeddings learnt with Web and Social Media data have competitive performances over supervised methods in the text based image retrieval task, and we clearly outperform state of the art in the MIRFlickr dataset when training in the target data. Further, we demonstrate how semantic multimodal image retrieval can be performed using the learnt embeddings, going beyond classical instance-level retrieval problems. Finally, we present a new dataset, InstaCities1M, composed by Instagram images and their associated texts that can be used for fair comparison of image-text embeddings.

📄 PDF Abstract BibTeX arXiv:1901.02004

Code (1)

gombru/LearnFromWebData 공식 구현

Tasks

Image RetrievalRetrievalSelf-Supervised Learning

Similar Papers 제목 키워드 기반

SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval

2026-05-15 · Wenjie Yang, Hang Yu, Yuyu Guo, Peng Di arxiv

In this work, we address the critical yet underexplored challenge of symmetric multimodal-to-multimodal (MM2MM) retrieval, where queries and contexts are interchangeable. Existing universal multimodal retrieval works str…

Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos

2021-04-26 · ICCV 2021 10 · Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne 외

Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this conte…

Action LocalizationClusteringContrastive LearningLong Video Retrieval (Background Removed)+5

ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval

2025-11-02 · Ahmed Masry, Megh Thakkar, Patrice Bechard, Sathwik Tejaswi Madhusudhan 외 arxiv

Retrieval-augmented generation has proven practical when models require specialized knowledge or access to the latest data. However, existing methods for multimodal document retrieval often replicate techniques developed…

Representation LearningContrastive Learning

Multimodal Whole Slide Foundation Model for Pathology

2024-11-29 · Tong Ding, Sophia J. Wagner, Andrew H. Song, Richard J. Chen 외

The field of computational pathology has been transformed with recent advances in foundation models that encode histopathology region-of-interests (ROIs) into versatile and transferable feature representations via self-s…

Cross-Modal RetrievalmodelPrognosisRetrieval+3

Contrastive Learning with Cross-Modal Knowledge Mining for Multimodal Human Activity Recognition

2022-05-20 · Razvan Brinzea, Bulat Khaertdinov, Stylianos Asteriadis

Human Activity Recognition is a field of research where input data can take many forms. Each of the possible input modalities describes human behaviour in a different way, and each has its own strengths and weaknesses. W…

Activity RecognitionContrastive LearningHuman Activity RecognitionRetrieval+1