A survey on Self Supervised learning approaches for improving Multimodal representation learning
Recently self supervised learning has seen explosive growth and use in variety of machine learning tasks because of its ability to avoid the cost of annotating large-scale datasets. This paper gives an overview for best self supervised learning approaches for multimodal learning. The presented approaches have been aggregated by extensive study of the literature and tackle the application of self supervised learning in different ways. The approaches discussed are cross modal generation, cross modal pretraining, cyclic translation, and generating unimodal labels in self supervised fashion.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation LearningSelf-Supervised LearningTranslationSimilar Papers 제목 키워드 기반
Self-Supervised Multimodal Learning: A Survey
Multimodal learning, which aims to understand and analyze information from multiple modalities, has achieved substantial progress in the supervised regime in recent years. However, the heavy dependence on data paired wit…
Machine TranslationSelf-Supervised LearningSurveySurvey on Self-Supervised Multimodal Representation Learning and Foundation Models
Deep learning has been the subject of growing interest in recent years. Specifically, a specific type called Multimodal learning has shown great promise for solving a wide range of problems in domains such as language, v…
Representation LearningSelf-Supervised LearningSurveySelf-Supervised Learning for Videos: A Survey
The remarkable success of deep learning in various domains relies on the availability of large-scale annotated datasets. However, obtaining annotations is expensive and requires great effort, which is especially challeng…
Contrastive LearningDomain GeneralizationRepresentation LearningSelf-Supervised Learning+1Multimodal Conversational AI: A Survey of Datasets and Approaches
As humans, we experience the world with all our senses or modalities (sound, sight, touch, smell, and taste). We use these modalities, particularly sight and touch, to convey and interpret specific meanings. Multimodal e…
SurveyTranslationA Survey of Self-Supervised and Few-Shot Object Detection
Labeling data is often expensive and time-consuming, especially for tasks such as object detection and instance segmentation, which require dense labeling of the image. While few-shot object detection is about training a…
Few-Shot Object DetectionInstance SegmentationObjectobject-detection+3