Order-Embeddings of Images and Language
Hypernymy, textual entailment, and image captioning can be seen as special cases of a single visual-semantic hierarchy over words, sentences, and images. In this paper we advocate for explicitly modeling the partial order structure of this hierarchy. Towards this goal, we introduce a general method for learning ordered representations, and show how it can be applied to a variety of tasks involving images and language. We show that the resulting representations improve performance over current approaches for hypernym prediction and image-caption retrieval.
Code (2)
Tasks
Cross-Modal RetrievalImage CaptioningNatural Language InferenceRetrievalSimilar Papers 제목 키워드 기반
UAEM-ITAM at SemEval-2022 Task 5: Vision-Language Approach to Recognize Misogynous Content in Memes
In the context of the Multimedia Automatic Misogyny Identification (MAMI) competition 2022, we developed a framework for extracting lexical-semantic features from text and combine them with semantic descriptions of image…
Dimensionality ReductionOrder embeddings and character-level convolutions for multimodal alignment
With the novel and fast advances in the area of deep neural networks, several challenging image-based tasks have been recently approached by researchers in pattern recognition and computer vision. In this paper, we addre…
RetrievalSemantic correspondenceWord EmbeddingsBridging Languages through Images with Deep Partial Canonical Correlation Analysis
We present a deep neural network that leverages images to improve bilingual text embeddings. Relying on bilingual image tags and descriptions, our approach conditions text embedding induction on the shared visual informa…
Image DescriptionImage RetrievalQuestion AnsweringRepresentation Learning+3Visual Grounding of Inter-lingual Word-Embeddings
Visual grounding of Language aims at enriching textual representations of language with multiple sources of visual knowledge such as images and videos. Although visual grounding is an area of intense research, inter-ling…
Visual GroundingWord EmbeddingsWord SimilarityRetrieving Similar E-Commerce Images Using Deep Learning
In this paper, we propose a deep convolutional neural network for learning the embeddings of images in order to capture the notion of visual similarity. We present a deep siamese architecture that when trained on positiv…
Deep LearningFine-Grained Visual RecognitionImage RetrievalProduct Recommendation+2