From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
We propose to use the visual denotations of linguistic expressions (i.e. the set of images they describe) to define novel denotational similarity metrics, which we show to be at least as beneficial as distributional similarities for two tasks that require semantic inference. To compute these denotational similarities, we construct a denotation graph, i.e. a subsumption hierarchy over constituents and their denotations, based on a large corpus of 30K images and 150K descriptive captions.
Code (0)
등록된 구현이 없습니다.
Tasks
DescriptiveSemantic Textual SimilaritySimilar Papers 제목 키워드 기반
Visual Denotations for Recognizing Textual Entailment
In the logic approach to Recognizing Textual Entailment, identifying phrase-to-phrase semantic relations is still an unsolved problem. Resources such as the Paraphrase Database offer limited coverage despite their large …
Natural Language InferenceSemantic CompositionImproving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
Brain-to-Image reconstruction aims to recover visual stimuli perceived by humans from brain activity. However, the reconstructed visual stimuli often missing details and semantic inconsistencies, which may be attributed …
Image ReconstructionLanguage ModelingLanguage ModellingSemantic Similarity+1VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions
We address the task of evaluating image description generation systems. We propose a novel image-aware metric for this task: VIFIDEL. It estimates the faithfulness of a generated caption with respect to the content of th…
Image DescriptionSemantic SimilaritySemantic Textual SimilarityLearning to Generate Compositional Color Descriptions
The production of color language is essential for grounded language generation. Color descriptions have many challenging properties: they can be vague, compositionally complex, and denotationally rich. We present an effe…
Language ModelingLanguage ModellingText GenerationZEST: Zero-shot Learning from Text Descriptions using Textual Similarity and Visual Summarization
We study the problem of recognizing visual entities from the textual descriptions of their classes. Specifically, given birds' images with free-text descriptions of their species, we learn to classify images of previousl…
Zero-Shot Learning