paper-with-me

Papers

On the Comparison between Multi-modal and Single-modal Contrastive Learning

2024-11-05 · Wei Huang, Andi Han, Yongqiang Chen, Yuan Cao, Zhiqiang Xu, Taiji Suzuki

Multi-modal contrastive learning with language supervision has presented a paradigm shift in modern machine learning. By pre-training on a web-scale dataset, multi-modal contrastive learning can learn high-quality representations that exhibit impressive robustness and transferability. Despite its empirical success, the theoretical understanding is still in its infancy, especially regarding its comparison with single-modal contrastive learning. In this work, we introduce a feature learning theory framework that provides a theoretical foundation for understanding the differences between multi-modal and single-modal contrastive learning. Based on a data generation model consisting of signal and noise, our analysis is performed on a ReLU network trained with the InfoMax objective function. Through a trajectory-based optimization analysis and generalization characterization on downstream tasks, we identify the critical factor, which is the signal-to-noise ratio (SNR), that impacts the generalizability in downstream tasks of both multi-modal and single-modal contrastive learning. Through the cooperation between the two modalities, multi-modal learning can achieve better feature learning, leading to improvements in performance in downstream tasks compared to single-modal learning. Our analysis provides a unified framework that can characterize the optimization and generalization of both single-modal and multi-modal contrastive learning. Empirical experiments on both synthetic and real-world datasets further consolidate our theoretical findings.

📄 PDF Abstract BibTeX arXiv:2411.02837

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningLearning Theory

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Multi-Format Contrastive Learning of Audio Representations

2021-03-11 · Luyu Wang, Aaron van den Oord

Recent advances suggest the advantage of multi-modal training in comparison with single-modal methods. In contrast to this view, in our work we find that similar gain can be obtained from training with different formats …

Audio ClassificationContrastive Learning

Generalized Semantic Preserving Hashing for N-Label Cross-Modal Retrieval

2017-07-01 · CVPR 2017 7 · Devraj Mandal, Kunal. N. Chaudhury, Soma Biswas

Due to availability of large amounts of multimedia data, cross-modal matching is gaining increasing importance. Hashing based techniques provide an attractive solution to this problem when the data size is large. Differe…

Cross-Modal RetrievalRetrievalSemantic SimilaritySemantic Textual Similarity

Application of Multimodal Fusion Deep Learning Model in Disease Recognition

2024-05-22 · Xiaoyi Liu, Hongjie Qiu, Muqing Li, Zhou Yu 외

This paper introduces an innovative multi-modal fusion deep learning approach to overcome the drawbacks of traditional single-modal recognition techniques. These drawbacks include incomplete information and limited diagn…

Deep LearningDiagnostic

Learning Multi-Modal Nonlinear Embeddings: Performance Bounds and an Algorithm

2020-06-03 · Semih Kaya, Elif Vural

While many approaches exist in the literature to learn low-dimensional representations for data collections in multiple modalities, the generalizability of multi-modal nonlinear embeddings to previously unseen data is a …

cross-modal alignmentGeneral Classificationimage-classificationImage Classification+5

Multimodal Sparse Coding for Event Detection

2016-05-17 · Youngjune Gwon, William Campbell, Kevin Brady, Douglas Sturim 외

Unsupervised feature learning methods have proven effective for classification tasks based on a single modality. We present multimodal sparse coding for learning feature representations shared across multiple modalities.…

ClassificationEvent DetectionGeneral Classification