paper-with-me

Papers

Visual Graphs from Motion (VGfM): Scene understanding with object geometry reasoning

2018-07-16 · Paul Gay, Stuart James, Alessio Del Bue

Recent approaches on visual scene understanding attempt to build a scene graph -- a computational representation of objects and their pairwise relationships. Such rich semantic representation is very appealing, yet difficult to obtain from a single image, especially when considering complex spatial arrangements in the scene. Differently, an image sequence conveys useful information using the multi-view geometric relations arising from camera motion. Indeed, in such cases, object relationships are naturally related to the 3D scene structure. To this end, this paper proposes a system that first computes the geometrical location of objects in a generic scene and then efficiently constructs scene graphs from video by embedding such geometrical reasoning. Such compelling representation is obtained using a new model where geometric and visual features are merged using an RNN framework. We report results on a dataset we created for the task of 3D scene graph generation in multiple views.

📄 PDF Abstract BibTeX arXiv:1807.05933

Code (2)

paulgay/VGfM 공식 구현 tf
ShunChengWu/3DSSG pytorch

Tasks

3d scene graph generationGraph GenerationScene Graph GenerationScene Understanding

Similar Papers 제목 키워드 기반

LibraGen: Playing a Balance Game in Subject-Driven Video Generation

2026-03-13 · Jiahao Zhu, Shanshan Lao, Lijie Liu, Gen Li 외 arxiv

With the advancement of video generation foundation models (VGFMs), customized generation, particularly subject-to-video (S2V), has attracted growing attention. However, a key challenge lies in balancing the intrinsic pr…

Video Generation

Joint Velocity-Growth Flow Matching for Single-Cell Dynamics Modeling

2025-05-19 · Dongyi Wang, Yuanwei Jiang, Zhenyi Zhang, Xiang Gu 외

Learning the underlying dynamics of single cells from snapshot data has gained increasing attention in scientific and machine learning research. The destructive measurement technique and cell proliferation/death result i…

NeuSyRE: Neuro-Symbolic Visual Understanding and Reasoning Framework based on Scene Graph Enrichment

2023-11-05 · Semantic Web 2023 11 · M. Jaleed Khan, John Breslin, Edward Curry

Neuro-symbolic hybrid approaches are inevitable for seamless high-level understanding and reasoning about visual scenes. Scene Graph Generation (SGG) is a symbolic image representation approach based on deep neural netwo…

Caption GenerationCommon Sense ReasoningGraph GenerationImage Captioning+7

Keep It CALM: Toward Calibration-Free Kilometer-Level SLAM with Visual Geometry Foundation Models via an Assistant Eye

2026-04-16 · Tianjun Zhang, Fengyi Zhang, Tianchen Deng, Lin Zhang 외 arxiv

Visual Geometry Foundation Models (VGFMs) demonstrate remarkable zero-shot capabilities in local reconstruction. However, deploying them for kilometer-level Simultaneous Localization and Mapping (SLAM) remains challengin…

Understanding the Role of Scene Graphs in Visual Question Answering

2021-01-14 · Vinay Damodaran, Sharanya Chakravarthy, Akshay Kumar, Anjana Umapathy 외

Visual Question Answering (VQA) is of tremendous interest to the research community with important applications such as aiding visually impaired users and image-based search. In this work, we explore the use of scene gra…

Graph GenerationQuestion AnsweringScene Graph GenerationVisual Question Answering+1