Visually Guided Spatial Relation Extraction from Text
Extraction of spatial relations from sentences with complex/nesting relationships is very challenging as often needs resolving inherent semantic ambiguities. We seek help from visual modality to fill the information gap in the text modality and resolve spatial semantic ambiguities. We use various recent vision and language datasets and techniques to train inter-modality alignment models, visual relationship classifiers and propose a novel global inference model to integrate these components into our structured output prediction model for spatial role and relation extraction. Our global inference model enables us to utilize the visual and geometric relationships between objects and improves the state-of-art results of spatial information extraction from text.
Code (0)
등록된 구현이 없습니다.
Tasks
Activity RecognitionImage CaptioningImage RetrievalObject LocalizationQuestion AnsweringRelationRelation ExtractionVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
RE$^2$: Region-Aware Relation Extraction from Visually Rich Documents
Current research in form understanding predominantly relies on large pre-trained language models, necessitating extensive data for pre-training. However, the importance of layout structure (i.e., the spatial relationship…
Graph AttentionRelationRelation ExtractionRelation PredictionMultimodal weighted graph representation for information extraction from visually rich documents.
This paper introduces a novel system for information extraction from visually rich documents (VRD) using a weighted graph representation. The proposed method aims to improve the performance of the information extraction …
Document Layout Analysisdocument understandingGraph Neural NetworkInformation Retrieval+2Towards Human-Like Machine Comprehension: Few-Shot Relational Learning in Visually-Rich Documents
Key-value relations are prevalent in Visually-Rich Documents (VRDs), often depicted in distinct spatial regions accompanied by specific color and font styles. These non-textual cues serve as important indicators that gre…
Document AIReading ComprehensionRelationRelational ReasoningA LayoutLMv3-Based Model for Enhanced Relation Extraction in Visually-Rich Documents
Document Understanding is an evolving field in Natural Language Processing (NLP). In particular, visual and spatial features are essential in addition to the raw text itself and hence, several multimodal models were deve…
document understandingKey Information ExtractionRelationRelation ExtractionGlobal Structure Knowledge-Guided Relation Extraction Method for Visually-Rich Document
Visual Relation Extraction (VRE) is a powerful means of discovering relationships between entities within visually-rich documents. Existing methods often focus on manipulating entity features to find pairwise relations, …
RelationRelation Extraction