paper-with-me

Papers

On the Generalization of Multi-modal Contrastive Learning

2023-06-07 · Qi Zhang, Yifei Wang, Yisen Wang

Multi-modal contrastive learning (MMCL) has recently garnered considerable interest due to its superior performance in visual tasks, achieved by embedding multi-modal data, such as visual-language pairs. However, there still lack theoretical understandings of how MMCL extracts useful visual representation from multi-modal pairs, and particularly, how MMCL outperforms previous approaches like self-supervised contrastive learning (SSCL). In this paper, by drawing an intrinsic connection between MMCL and asymmetric matrix factorization, we establish the first generalization guarantees of MMCL for visual downstream tasks. Based on this framework, we further unify MMCL and SSCL by showing that MMCL implicitly performs SSCL with (pseudo) positive pairs induced by text pairs. Through this unified perspective, we characterize the advantage of MMCL by showing that text pairs induce more semantically consistent and diverse positive pairs, which, according to our analysis, provably benefit downstream generalization. Inspired by this finding, we propose CLIP-guided resampling methods to significantly improve the downstream performance of SSCL on ImageNet by leveraging multi-modal information. Code is available at https://github.com/PKU-ML/CLIP-Help-SimCLR.

📄 PDF Abstract BibTeX arXiv:2306.04272

Code (1)

pku-ml/clip-help-simclr 공식 구현 pytorch

Tasks

Contrastive Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

On the Comparison between Multi-modal and Single-modal Contrastive Learning

2024-11-05 · Wei Huang, Andi Han, Yongqiang Chen, Yuan Cao 외

Multi-modal contrastive learning with language supervision has presented a paradigm shift in modern machine learning. By pre-training on a web-scale dataset, multi-modal contrastive learning can learn high-quality repres…

Contrastive LearningLearning Theory

A Statistical Theory of Contrastive Pre-training and Multimodal Generative AI

2025-01-08 · Kazusato Oko, Licong Lin, Yuhang Cai, Song Mei

Multi-modal generative AI systems, such as those combining vision and language, rely on contrastive pre-training to learn representations across different modalities. While their practical benefits are widely acknowledge…

zero-shot-classificationZero-Shot Learning

Hybrid Contrastive Learning of Tri-Modal Representation for Multimodal Sentiment Analysis

2021-09-04 · Sijie Mai, Ying Zeng, Shuangjia Zheng, Haifeng Hu

The wide application of smart devices enables the availability of multimodal data, which can be utilized in many tasks. In the field of multimodal sentiment analysis (MSA), most previous works focus on exploring intra- a…

Contrastive LearningMultimodal Sentiment AnalysisSentiment Analysis

Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning

2026-03-20 · Jianan Huang, Rodolfo V. Valentim, Luca Vassio, Matteo Boffa 외 arxiv

The use of ML in cybersecurity has long been impaired by generalization issues: Models that work well in controlled scenarios fail to maintain performance in production. The root cause often lies in ML algorithms learnin…

Contrastive Learning

Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination

2026-08-20 · Kaixin Xu, NaiJin Liu, Yulin Kang, Tangyue Jin 외 arxiv

Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data is missing at random or are designed specifically for certain missing patterns, …

Multimodal Sentiment AnalysisContrastive Learning