Cross-Modal Contrastive Representation Learning for Audio-to-Image Generation
Multiple modalities for certain information provide a variety of perspectives on that information, which can improve the understanding of the information. Thus, it may be crucial to generate data of different modality from the existing data to enhance the understanding. In this paper, we investigate the cross-modal audio-to-image generation problem and propose Cross-Modal Contrastive Representation Learning (CMCRL) to extract useful features from audios and use it in the generation phase. Experimental results show that CMCRL enhances quality of images generated than previous research.
Code (0)
등록된 구현이 없습니다.
Tasks
Image GenerationRepresentation LearningSimilar Papers 제목 키워드 기반
Distilling Audio-Visual Knowledge by Compositional Contrastive Learning
Having access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to transfer knowledge across heterogeneous m…
Audio Taggingaudio-visual learningContrastive LearningKnowledge Distillation+3Connecting Multi-modal Contrastive Representations
Multi-modal Contrastive Representation learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across vari…
3D Point Cloud ClassificationcounterfactualImage RetrievalPoint Cloud Classification+2Improving Sound Source Localization with Joint Slot Attention on Image and Audio
Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embed…
Contrastive LearningCross-Modal RetrievalSound Source LocalizationExtending Multi-modal Contrastive Representations
Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality p…
3D Object ClassificationRepresentation LearningText RetrievalOn the Language Encoder of Contrastive Cross-modal Models
Contrastive cross-modal models such as CLIP and CLAP aid various vision-language (VL) and audio-language (AL) tasks. However, there has been limited investigation of and improvement in their language encoder, which is th…
cross-modal alignmentSentenceSentence EmbeddingSentence-Embedding