paper-with-me

홈 › Papers

Self-supervised visual learning in the low-data regime: a comparative evaluation

2024-04-26 · Sotirios Konstantakos, Jorgen Cani, Ioannis Mademlis, Despina Ioanna Chalkiadaki, Yuki M. Asano, Efstratios Gavves, Georgios Th. Papadopoulos

Self-Supervised Learning (SSL) is a valuable and robust training methodology for contemporary Deep Neural Networks (DNNs), enabling unsupervised pretraining on a 'pretext task' that does not require ground-truth labels/annotation. This allows efficient representation learning from massive amounts of unlabeled training data, which in turn leads to increased accuracy in a 'downstream task' by exploiting supervised transfer learning. Despite the relatively straightforward conceptualization and applicability of SSL, it is not always feasible to collect and/or to utilize very large pretraining datasets, especially when it comes to real-world application settings. In particular, in cases of specialized and domain-specific application scenarios, it may not be achievable or practical to assemble a relevant image pretraining dataset in the order of millions of instances or it could be computationally infeasible to pretrain at this scale, e.g., due to unavailability of sufficient computational resources that SSL methods typically require to produce improved visual analysis results. This situation motivates an investigation on the effectiveness of common SSL pretext tasks, when the pretraining dataset is of relatively limited/constrained size. This work briefly introduces the main families of modern visual SSL methods and, subsequently, conducts a thorough comparative experimental evaluation in the low-data regime, targeting to identify: a) what is learnt via low-data SSL pretraining, and b) how do different SSL categories behave in such training scenarios. Interestingly, for domain-specific downstream tasks, in-domain low-data SSL pretraining outperforms the common approach of large-scale pretraining on general datasets.

📄 PDF Abstract BibTeX arXiv:2404.17202

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSelf-Supervised LearningTransfer Learning

Similar Papers 제목 키워드 기반

Audio-Visual Speech Enhancement and Separation by Utilizing Multi-Modal Self-Supervised Embeddings

2022-10-31 · I-Chun Chern, Kuo-Hsuan Hung, Yi-Ting Chen, Tassadaq Hussain 외

AV-HuBERT, a multi-modal self-supervised learning model, has been shown to be effective for categorical problems such as automatic speech recognition and lip-reading. This suggests that useful audio-visual speech represe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Lip Readingregression+5

Trading robust representations for sample complexity through self-supervised visual experience

2018-12-01 · NeurIPS 2018 12 · Andrea Tacchetti, Stephen Voinea, Georgios Evangelopoulos

Learning in small sample regimes is among the most remarkable features of the human perceptual system. This ability is related to robustness to transformations, which is acquired through visual experience in the form of …

One-Shot LearningRepresentation LearningSelf-Supervised Learning

LiRA: Learning Visual Speech Representations from Audio through Self-supervision

2021-06-16 · Pingchuan Ma, Rodrigo Mira, Stavros Petridis, Björn W. Schuller 외

The large amount of audiovisual content being shared online today has drawn substantial attention to the prospect of audiovisual self-supervised learning. Recent works have focused on each of these modalities separately,…

Lip ReadingSelf-Supervised LearningSentence

Interpretable agent communication from scratch (with a generic visual processor emerging on the side)

2021-06-08 · NeurIPS 2021 12 · Roberto Dessì, Eugene Kharitonov, Marco Baroni

As deep networks begin to be deployed as autonomous agents, the issue of how they can communicate with each other becomes important. Here, we train two deep nets from scratch to perform realistic referent identification …

Self-Supervised Learning

DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning

2025-09-22 · ThankGod Egbe, Peng Wang, Zhihao Guo, Zidong Chen arxiv

This paper evaluates DINOv3, a recent large-scale self-supervised vision backbone, for visuomotor diffusion policy learning in robotic manipulation. We investigate whether a purely self-supervised encoder can match or su…