Spatial Language Understanding with Multimodal Graphs using Declarative Learning based Programming
This work is on a previously formalized semantic evaluation task of spatial role labeling (SpRL) that aims at extraction of formal spatial meaning from text. Here, we report the results of initial efforts towards exploiting visual information in the form of images to help spatial language understanding. We discuss the way of designing new models in the framework of declarative learning-based programming (DeLBP). The DeLBP framework facilitates combining modalities and representing various data in a unified graph. The learning and inference models exploit the structure of the unified graph as well as the global first order domain constraints beyond the data to predict the semantics which forms a structured meaning representation of the spatial context. Continuous representations are used to relate the various elements of the graph originating from different modalities. We improved over the state-of-the-art results on SpRL.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningImage RetrievalQuestion AnsweringStructured PredictionVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Spatial Symmetry Driven Pruning Strategies for Efficient Declarative Spatial Reasoning
Declarative spatial reasoning denotes the ability to (declaratively) specify and solve real-world problems related to geometric and qualitative spatial representation and reasoning within standard knowledge representatio…
Spatial ReasoningBoosting Audio Visual Question Answering via Key Semantic-Aware Cues
The Audio Visual Question Answering (AVQA) task aims to answer questions related to various visual objects, sounds, and their interactions in videos. Such naturally multimodal videos contain rich and complex dynamic audi…
Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Question AnsweringVisual Question AnsweringBridging Modalities and Transferring Knowledge: Enhanced Multimodal Understanding and Recognition
This manuscript explores multimodal alignment, translation, fusion, and transference to enhance machine understanding of complex inputs. We organize the work into five chapters, each addressing unique challenges in multi…
Knowledge DistillationAction RecognitionObject DetectionKnowledge GraphsEchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
Recent progress in Vision-Language-Action (VLA) models has enabled embodied agents to interpret multimodal instructions and perform complex tasks. However, existing VLAs are mostly confined to short-horizon, table-top ma…
Talking about the Moving Image: A Declarative Model for Image Schema Based Embodied Perception Grounding and Language Generation
We present a general theory and corresponding declarative model for the embodied grounding and natural language based analytical summarisation of dynamic visuo-spatial imagery. The declarative model ---ecompassing spatio…
Spatial ReasoningText Generation