Learning Cross-Modal Deep Embeddings for Multi-Object Image Retrieval using Text and Sketch
In this work we introduce a cross modal image retrieval system that allows both text and sketch as input modalities for the query. A cross-modal deep network architecture is formulated to jointly model the sketch and text input modalities as well as the the image output modality, learning a common embedding between text and images and between sketches and images. In addition, an attention model is used to selectively focus the attention on the different objects of the image, allowing for retrieval with multiple objects in the query. Experiments show that the proposed method performs the best in both single and multiple object image retrieval in standard datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Image RetrievalRetrievalSimilar Papers 제목 키워드 기반
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task
In this paper, we propose a new approach to learn multimodal multilingual embeddings for matching images and their relevant captions in two languages. We combine two existing objective functions to make images and captio…
Cross-Modal RetrievalImage to textImage-to-Text RetrievalMultilingual Word Embeddings+3Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
Multimodal sentence embedding models typically leverage image-caption pairs in addition to textual data during training. However, such pairs often contain noise, including redundant or irrelevant information on either th…
Semantic Textual SimilarityRepresentation LearningContrastive LearningObject DetectionMultimodal Skip-gram Using Convolutional Pseudowords
This work studies the representational mapping across multimodal data such that given a piece of the raw data in one modality the corresponding semantic description in terms of the raw data in another modality is immedia…
Object RecognitionRetrievalVideo RetrievalWord SimilarityModalChorus: Visual Probing and Alignment of Multi-modal Embeddings via Modal Fusion Map
Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal f…
Cross-Modal RetrievalDimensionality Reductionzero-shot-classificationZero-Shot LearningObject-X: Learning to Reconstruct Multi-Modal 3D Object Representations
Learning effective multi-modal 3D representations of objects is essential for numerous applications, such as augmented reality and robotics. Existing methods often rely on task-specific embeddings that are tailored eithe…
3D Object ReconstructionNovel View SynthesisObjectObject Reconstruction