paper-with-me

홈 › Papers

Understanding Self-supervised Learning with Dual Deep Networks

2020-10-01 · Yuandong Tian, Lantao Yu, Xinlei Chen, Surya Ganguli

We propose a novel theoretical framework to understand contrastive self-supervised learning (SSL) methods that employ dual pairs of deep ReLU networks (e.g., SimCLR). First, we prove that in each SGD update of SimCLR with various loss functions, including simple contrastive loss, soft Triplet loss and InfoNCE loss, the weights at each layer are updated by a \emph{covariance operator} that specifically amplifies initial random selectivities that vary across data samples but survive averages over data augmentations. To further study what role the covariance operator plays and which features are learned in such a process, we model data generation and augmentation processes through a \emph{hierarchical latent tree model} (HLTM) and prove that the hidden neurons of deep ReLU networks can learn the latent variables in HLTM, despite the fact that the network receives \emph{no direct supervision} from these unobserved latent variables. This leads to a provable emergence of hierarchical features through the amplification of initially random selectivities through contrastive SSL. Extensive numerical studies justify our theoretical findings. Code is released in https://github.com/facebookresearch/luckmatters/tree/master/ssl.

📄 PDF Abstract BibTeX arXiv:2010.00578

Code (2)

facebookresearch/luckmatters 공식 구현 pytorch
facebookresearch/luckmatters/tree/master/ssl 공식 구현 pytorch

Tasks

Self-Supervised LearningTriplet

Methods 이 논문이 사용한 방법론

Triplet Loss The goal of Triplet loss, in the context of Siamese Networks, is to maximize the joint probability among all score-pairs i.e. the product of all probabilities. By using its…
InfoNCE 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Batch Normalization 설명 없음
Residual Connection 설명 없음
Residual Block Residual Blocks are skip-connection blocks that learn residual functions with reference to the layer inputs, instead of learning unreferenced functions. They were introduced…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Bottleneck Residual Block A Bottleneck Residual Block is a variant of the residual block that utilises 1x1 convolutions to create a bottleneck. The…

Similar Papers 제목 키워드 기반

MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

2025-01-02 · Haina Zhu, Yizhi Zhou, Hangting Chen, Jianwei Yu 외

Recent years have witnessed the success of foundation models pre-trained with self-supervised learning (SSL) in various music informatics understanding tasks, including music tagging, instrument classification, key detec…

Contrastive LearningKey DetectionMusic TaggingQuantization+2

Reinforcing Multimodal Understanding and Generation with Dual Self-rewards

2025-06-09 · Jixiang Hong, Yiran Zhang, Guanzhong Wang, Yi Liu 외

Building upon large language models (LLMs), recent large multimodal models (LMMs) unify cross-model understanding and generation into a single framework. However, LMMs still struggle to achieve accurate image-text alignm…

Multiplexed Immunofluorescence Brain Image Analysis Using Self-Supervised Dual-Loss Adaptive Masked Autoencoder

2022-05-10 · Son T. Ly, Bai Lin, Hung Q. Vo, Dragan Maric 외

Reliable large-scale cell detection and segmentation is the fundamental first step to understanding biological processes in the brain. The ability to phenotype cells at scale can accelerate preclinical drug evaluation an…

Cell DetectionContrastive LearningImage ReconstructionSegmentation+2

Self-Supervised Object Detection from Egocentric Videos

2023-01-01 · ICCV 2023 1 · Peri Akiva, Jing Huang, Kevin J Liang, Rama Kovvuri 외

Understanding the visual world from the perspective of humans (egocentric) has been a long-standing challenge in computer vision. Egocentric videos exhibit high scene complexity and irregular motion flows compared to…

Class-agnostic Object DetectionObjectobject-detectionObject Detection+2

MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

2023-05-31 · Yizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma 외

Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speec…

Language ModellingQuantizationSelf-Supervised Learning