Multimodal Representation Learning via Maximization of Local Mutual Information
We propose and demonstrate a representation learning approach by maximizing the mutual information between local features of images and text. The goal of this approach is to learn useful image representations by taking advantage of the rich information contained in the free text that describes the findings in the image. Our method trains image and text encoders by encouraging the resulting representations to exhibit high local mutual information. We make use of recent advances in mutual information estimation with neural network discriminators. We argue that the sum of local mutual information is typically a lower bound on the global mutual information. Our experimental results in the downstream image classification tasks demonstrate the advantages of using local features for image-text representation learning.
Code (1)
Tasks
image-classificationImage ClassificationMutual Information EstimationRepresentation LearningSimilar Papers 제목 키워드 기반
Multimodal Representations Learning Based on Mutual Information Maximization and Minimization and Identity Embedding for Multimodal Sentiment Analysis
Multimodal sentiment analysis (MSA) is a fundamental complex research problem due to the heterogeneity gap between different modalities and the ambiguity of human emotional expression. Although there have been many succe…
Multimodal Sentiment AnalysisSentiment AnalysisInfoSeg: Unsupervised Semantic Image Segmentation with Mutual Information Maximization
We propose a novel method for unsupervised semantic image segmentation based on mutual information maximization between local and global high-level image features. The core idea of our work is to leverage recent progress…
Image SegmentationRepresentation LearningSemantic SegmentationUnsupervised Semantic SegmentationEnhanced Multimodal Representation Learning with Cross-modal KD
This paper explores the tasks of leveraging auxiliary modalities which are only available at training to enhance multimodal representation learning through cross-modal Knowledge Distillation (KD). The widely adopted mutu…
Contrastive LearningEmotion ClassificationKnowledge DistillationRepresentation Learning+3Self-MI: Efficient Multimodal Fusion via Self-Supervised Multi-Task Learning with Auxiliary Mutual Information Maximization
Multimodal representation learning poses significant challenges in capturing informative and distinct features from multiple modalities. Existing methods often struggle to exploit the unique characteristics of each modal…
Multi-Task LearningRepresentation LearningSelf-Supervised LearningFast computation of mutual information in the frequency domain with applications to global multimodal image alignment
Multimodal image alignment is the process of finding spatial correspondences between images formed by different imaging techniques or under different conditions, to facilitate heterogeneous data fusion and correlative an…
GPU