Image Annotation with ISO-Space: Distinguishing Content from Structure
Natural language descriptions of visual media present interesting problems for linguistic annotation of spatial information. This paper explores the use of ISO-Space, an annotation specification to capturing spatial information, for encoding spatial relations mentioned in descriptions of images. Especially, we focus on the distinction between references to representational content and structural components of images, and the utility of such a distinction within a compositional semantics. We also discuss how such a structure-content distinction within the linguistic annotation can be leveraged to compute further inferences about spatial configurations depicted by images with verbal captions. We construct a composition table to relate content-based relations to structure-based relations in the image, as expressed in the captions. While still preliminary, our initial results suggest that a weak composition table is both sound and informative for deriving new spatial relations.
Code (0)
등록된 구현이 없습니다.
Tasks
Content-Based Image RetrievalImage CaptioningImage RetrievalSemantic Role LabelingSimilar Papers 제목 키워드 기반
Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion
Transformer-based multimodal models are widely used in industrial-scale recommendation, search, and advertising systems for content understanding and relevance ranking. Enhancing labeled training data quality and cross-m…
Active LearningLanguage ModelingLanguage ModellingDistinguishing Enzyme Structures from Non-enzymes Without Alignments
The ability to predict protein function from structure is becoming increasingly important as the number of structures resolved is growing more rapidly than our capacity to study function. Current methods for predicting p…
Graph ClassificationFrom Pixel to Slide image: Polarization Modality-based Pathological Diagnosis Using Representation Learning
Thyroid cancer is the most common endocrine malignancy, and accurately distinguishing between benign and malignant thyroid tumors is crucial for developing effective treatment plans in clinical practice. Pathologically, …
DiagnosticRepresentation LearningEvaluating the Utility of Grounding Documents with Reference-Free LLM-based Metrics
Retrieval Augmented Generation (RAG)'s success depends on the utility the LLM derives from the content used for grounding. Quantifying content utility does not have a definitive specification and existing metrics ignore …
Hash Function Learning via Codewords
In this paper we introduce a novel hash learning framework that has two main distinguishing features, when compared to past approaches. First, it utilizes codewords in the Hamming space as ancillary means to accomplish i…
Content-Based Image RetrievalImage RetrievalRetrieval