Learning Semantic Relationship Among Instances for Image-Text Matching
Image-text matching, a bridge connecting image and language, is an important task, which generally learns a holistic cross-modal embedding to achieve a high-quality semantic alignment between the two modalities. However, previous studies only focus on capturing fragment-level relation within a sample from a particular modality, e.g., salient regions in an image or text words in a sentence, where they usually pay less attention to capturing instance-level interactions among samples and modalities, e.g., multiple images and texts. In this paper, we argue that sample relations could help learn subtle differences for hard negative instances, and thus transfer shared knowledge for infrequent samples should be promising in obtaining better holistic embeddings. Therefore, we propose a novel hierarchical relation modeling framework (HREM), which explicitly capture both fragment- and instance-level relations to learn discriminative and robust cross-modal embeddings. Extensive experiments on Flickr30K and MS-COCO show our proposed method outperforms the state-of-the-art ones by 4%-10% in terms of rSum.
Code (1)
Tasks
Cross-Modal RetrievalImage RetrievalImage-text matchingMultimodal Deep LearningNetwork EmbeddingRelationRetrievalSentenceText MatchingText Retrievaltext similaritySimilar Papers 제목 키워드 기반
Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
Image-text matching is crucial for bridging the semantic gap between computer vision and natural language processing. However, existing methods still face challenges in handling high-order associations and semantic ambig…
Contrastive LearningImage-text matchingContrastive Learning Subspace for Text Clustering
Contrastive learning has been frequently investigated to learn effective representations for text clustering tasks. While existing contrastive learning-based text clustering methods only focus on modeling instance-wise s…
ClusteringContrastive LearningSemantic SimilaritySemantic Textual Similarity+1Multi-Instance Learning by Utilizing Structural Relationship among Instances
Multi-Instance Learning(MIL) aims to learn the mapping between a bag of instances and the bag-level label. Therefore, the relationships among instances are very important for learning the mapping. In this paper, we propo…
Graph Attentionimage-classificationImage ClassificationMedical Image ClassificationLearning Structured Semantic Embeddings for Visual Recognition
Numerous embedding models have been recently explored to incorporate semantic knowledge into visual recognition. Existing methods typically focus on minimizing the distance between the corresponding images and texts in t…
General ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONWord Embeddings+1Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
Despite progress in visual perception tasks such as image classification and detection, computers still struggle to understand the interdependency of objects in the scene as a whole, e.g., relations between objects or th…
Attributeimage-classificationImage ClassificationObject+5