paper-with-me

Papers

Cross-Modal Contrastive Representation Learning for Audio-to-Image Generation

2022-07-20 · HaeChun Chung, JooYong Shim, Jong-Kook Kim

Multiple modalities for certain information provide a variety of perspectives on that information, which can improve the understanding of the information. Thus, it may be crucial to generate data of different modality from the existing data to enhance the understanding. In this paper, we investigate the cross-modal audio-to-image generation problem and propose Cross-Modal Contrastive Representation Learning (CMCRL) to extract useful features from audios and use it in the generation phase. Experimental results show that CMCRL enhances quality of images generated than previous research.

📄 PDF Abstract BibTeX arXiv:2207.12121

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationRepresentation Learning

Similar Papers 제목 키워드 기반

Distilling Audio-Visual Knowledge by Compositional Contrastive Learning

2021-04-22 · CVPR 2021 1 · Yanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan 외

Having access to multi-modal cues (e.g. vision and audio) empowers some cognitive tasks to be done faster compared to learning from a single modality. In this work, we propose to transfer knowledge across heterogeneous m…

Audio Taggingaudio-visual learningContrastive LearningKnowledge Distillation+3

Connecting Multi-modal Contrastive Representations

2023-05-22 · NeurIPS 2023 11 · Zehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang 외

Multi-modal Contrastive Representation learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across vari…

3D Point Cloud ClassificationcounterfactualImage RetrievalPoint Cloud Classification+2

Improving Sound Source Localization with Joint Slot Attention on Image and Audio

2025-04-21 · CVPR 2025 1 · Inho Kim, Youngkil Song, Jicheol Park, Won Hwa Kim 외

Sound source localization (SSL) is the task of locating the source of sound within an image. Due to the lack of localization labels, the de facto standard in SSL has been to represent an image and audio as a single embed…

Contrastive LearningCross-Modal RetrievalSound Source Localization

Extending Multi-modal Contrastive Representations

2023-10-13 · Zehan Wang, Ziang Zhang, Luping Liu, Yang Zhao 외

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality p…

3D Object ClassificationRepresentation LearningText Retrieval

On the Language Encoder of Contrastive Cross-modal Models

2023-10-20 · Mengjie Zhao, Junya Ono, Zhi Zhong, Chieh-Hsin Lai 외

Contrastive cross-modal models such as CLIP and CLAP aid various vision-language (VL) and audio-language (AL) tasks. However, there has been limited investigation of and improvement in their language encoder, which is th…

cross-modal alignmentSentenceSentence EmbeddingSentence-Embedding