Visual Objects As Context: Exploiting Visual Objects for Lexical Entailment
We propose a new word representation method derived from visual objects in associated images to tackle the lexical entailment task. Although it has been shown that the \textit{Distributional Informativeness Hypothesis} (DIH) holds on text, in which the DIH assumes that a context surrounding a hyponym is more informative than that of a hypernym, it has never been tested on visual objects. Since our perception is tightly associated with language, it is meaningful to explore whether the DIH holds on visual objects. To this end, we consider visual objects as the context of a word and represent a word as a bag of visual objects found in images associated with the word. This allows us to test the feasibility of the visual DIH. To better distinguish word pairs in a hypernym relation from other relations such as co-hypernyms, we also propose a new measurable function that takes into account both the difference in the generality of meaning and similarity of meaning between words. Our experimental results show that the DIH holds on visual objects and that the proposed method combined with the proposed function outperforms existing unsupervised representation methods.
Code (0)
등록된 구현이 없습니다.
Tasks
InformativenessLexical EntailmentSimilar Papers 제목 키워드 기반
Exploiting Contextual Objects and Relations for 3D Visual Grounding
3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, thi…
Learning Object Semantic Similarity with Self-Supervision
Humans judge the similarity of two objects not just based on their visual appearance but also based on their semantic relatedness. However, it remains unclear how humans learn about semantic relationships between objects…
ObjectSemantic SimilaritySemantic Textual SimilarityTemporal SequencesMETOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection
Open-vocabulary video visual relationship detection aims to detect objects and their relationships in videos without being restricted by predefined object or relationship categories. Existing methods leverage the rich se…
Objectobject-detectionObject DetectionRelationship Detection+1Learning to Ground Visual Objects for Visual Dialog
Visual dialog is challenging since it needs to answer a series of coherent questions based on understanding the visual environment. How to ground related visual objects is one of the key problems. Previous studies utiliz…
Visual DialogTarget-Tailored Source-Transformation for Scene Graph Generation
Scene graph generation aims to provide a semantic and structural description of an image, denoting the objects (with nodes) and their relationships (with edges). The best performing works to date are based on exploiting …
graph constructionGraph GenerationObjectobject-detection+4