paper-with-me

홈 › Papers

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes

2025-05-06 · Sergey Linok, Vadim Semenov, Anastasia Trunova, Oleg Bulichev, Dmitry Yudin

The analysis of events in dynamic environments poses a fundamental challenge in the development of intelligent agents and robots capable of interacting with humans. Current approaches predominantly utilize visual models. However, these methods often capture information implicitly from images, lacking interpretable spatial-temporal object representations. To address this issue we introduce DyGEnc - a novel method for Encoding a Dynamic Graph. This method integrates compressed spatial-temporal structural observation representation with the cognitive capabilities of large language models. The purpose of this integration is to enable advanced question answering based on a sequence of textual scene graphs. Extended evaluations on the STAR and AGQA datasets indicate that DyGEnc outperforms existing visual methods by a large margin of 15-25% in addressing queries regarding the history of human-to-object interactions. Furthermore, the proposed method can be seamlessly extended to process raw input images utilizing foundational models for extracting explicit textual scene graphs, as substantiated by the results of a robotic experiment conducted with a wheeled manipulator platform. We hope that these findings will contribute to the implementation of robust and compressed graph-based robotic memory for long-horizon reasoning. Code is available at github.com/linukc/DyGEnc.

📄 PDF Abstract BibTeX arXiv:2505.03581

Code (1)

linukc/dygenc 공식 구현 pytorch

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Extend, don’t rebuild: Phrasing conditional graph modification as autoregressive sequence labelling

2021-11-01 · EMNLP 2021 11 · Leon Weber, Jannes Münchmeyer, Samuele Garda, Ulf Leser

Deriving and modifying graphs from natural language text has become a versatile basis technology for information extraction with applications in many subfields, such as semantic parsing or knowledge graph construction. A…

graph constructionGraph GenerationSemantic Parsing

SPAN: Learning Similarity between Scene Graphs and Images with Transformers

2023-04-02 · Yuren Cong, Wentong Liao, Bodo Rosenhahn, Michael Ying Yang

Learning similarity between scene graphs and images aims to estimate a similarity score given a scene graph and an image. There is currently no research dedicated to this task, although it is critical for scene graph gen…

Contrastive LearningGraph GenerationImage RetrievalRetrieval+2

SGRAM: Improving Scene Graph Parsing via Abstract Meaning Representation

2022-10-17 · Woo Suk Choi, Yu-Jung Heo, Byoung-Tak Zhang

Scene graph is structured semantic representation that can be modeled as a form of graph from images and texts. Image-based scene graph generation research has been actively conducted until recently, whereas text-based s…

Abstract Meaning RepresentationDependency ParsingGraph GenerationImage Retrieval+5

Scene Graph Parsing via Abstract Meaning Representation in Pre-trained Language Models

2022-07-01 · NAACL (DLG4NLP) 2022 7 · Woo Suk Choi, Yu-Jung Heo, Dharani Punithan, Byoung-Tak Zhang

In this work, we propose the application of abstract meaning representation (AMR) based semantic parsing models to parse textual descriptions of a visual scene into scene graphs, which is the first work to the best of ou…

Abstract Meaning RepresentationAMR ParsingDependency ParsingSemantic Parsing

Graph Similarities and Dual Approach for Sequential Text-to-Image Retrieval

2021-09-29 · Keonwoo Kim, Sihyeon Jo, Seong-Woo Kim

Sequential text-to-image retrieval, a.k.a. Story-to-images task, requires semantic alignment with a given story and maintaining global coherence in drawn image sequence simultaneously. Most of the previous works have onl…

Graph EmbeddingImage RetrievalRetrievalSentence+2