paper-with-me

Papers Scene Graph Generation

“Scene Graph Generation” 태그가 달린 논문 365편 · 필터 해제

Sterilizable Scene Graph Generation for Operating Rooms

2026-08-17 · Nick Lemke, Ssharvien Kumar Sivakumar, Antoine P. Sanner, John Kalkhof 외 arxiv

Scene graph generation from surgical video enables a holistic and structured understanding of surgical scenes by modeling objects and their semantic relationships. Despite recent advances, state-of-the-art approaches rel…

Scene Graph GenerationScene UnderstandingObject DetectionVideo Captioning

Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

2026-08-06 · Giorgio Tonetti, Laurent Kneip, Abel Gawel, Marco Hutter arxiv

Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or str…

Scene Graph GenerationSpatial Reasoning

Modeling Scientific Experiment Scenes: Dataset and Model

2026-08-03 · Minghao Zou, Qingtian Zeng, Shangkun Liu, Cong Liu 외 arxiv

Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily life images and overlook scientific experiment scenes with specialized instruments, task-speci…

Scene Graph GenerationRelational Reasoning

PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation

2026-07-08 · Yi Yang, Myrna Castillo, Bodo Rosenhahn, Michael Ying Yang arxiv

Online 3D scene graph generation builds a persistent, structured representation of a scene by incrementally fusing 2D observations into a global 3D graph. Existing online methods treat this fusion as a fully deterministi…

Scene Graph GenerationScene Understanding

Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability

2026-07-07 · Runfeng Qu, Pia K Bideau, Ole Hall, Julie Ouerfelli-Ethier 외 arxiv

Scene graph generation (SGG) approaches can be broadly classified into detector-based and query-based methods according to their underlying reasoning mechanisms. However, the discrepancy in their predictive behaviors, in…

Scene Graph Generation

NoPA: Non-Parametric Online 3D Scene Graph Generation

2026-07-01 · Qi Xun Yeo, Seungjun Lee, Yan Li, Gim Hee Lee hf

Classic 3D scene graph generation approaches fail to work in real-time due to the heavy computational cost of environment mapping and the need to generate intermediate point-cloud representations. To alleviate this issue…

Scene Graph GenerationPoint Clouds

DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors

2026-07-01 · Seok-Young Kim, Abdelrahman Elskhawy, Taewook Ha, Dooyoung Kim 외 arxiv

We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences. Existing methods often struggle to construct reliable 3D scene graphs due to unstable 3D obje…

Scene Graph Generation

OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments

2026-06-29 · Yirum Kim, Ue-Hwan Kim arxiv

3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments. Although advances in foundation models have enabled open-vocabulary 3DSG generation, existing approaches remain object-centric and …

Scene Graph GenerationRelational Reasoning

Not All Relations Rotate Alike: Transformation-Aware Decoupling for Viewpoint-Robust 3D Scene Graph Generation

2026-06-25 · Jingjun Sun, Chaowei Wang, Zhirui Liu, Jiaxu Tian 외 arxiv

3D Scene Graph Generation (3DSGG) represents 3D scenes as structured object-relation-object graphs, providing a compact relational abstraction for spatial understanding. In embodied intelligence settings, the same 3D sce…

Scene Graph Generation

Modular Diffusion Models for Structured Visual Recognition

2026-06-21 · Siddhesh Khandelwal, Björn Ommer, Leonid Sigal arxiv

Traditional supervised methods for structured visual recognition tasks -- such as object detection, segmentation, and scene graph generation -- often produce deterministic, fixed outputs, limiting their ability to captur…

Scene Graph GenerationInstance SegmentationObject Detection

SGFormer++: Semantic Graph Transformer for Incremental 3D Scene Graph Generation

2026-06-13 · Mengshi Qi, Changsheng Lv, Zijian Fu, Xianlin Zhang 외 arxiv

In this paper, we propose SGFormer++, a novel Semantic Graph Transformer for 3D scene graph generation (SGG), which aims to parse point cloud scenes into semantic structural graphs, where nodes denote detected object ins…

Scene Graph GenerationGraph Embedding

Visual Commonsense Driven Knowledge Refinements for Scene Graph Generation

2026-06-04 · Maëlic Neau, Salim Baloch, Jakob Suchan, Zoe Falomir 외 arxiv

Learning-driven Scene Graph Generation (SGG) models excel on frequent relation types but degrade sharply under annotation sparsity, failing to capture reliable visual commonsense knowledge. We propose a model-agnostic, s…

Visual Commonsense ReasoningScene Graph Generation

QPredSGG: Hybrid Quantum Predicate Learning for Long-Tailed Scene Graph Generation

2026-06-03 · Prerana Ramkumar, Nouhaila Innan, Muhammad Shafique arxiv

Scene Graph Generation (SGG) requires relational reasoning over objects and their interactions, but performance is often limited by severe long-tail predicate imbalance. Classical SGG models frequently rely on dataset st…

Scene Graph GenerationRelational ReasoningVisual Reasoning

Seeing Fast and Slow: Bimodal 3D Scene Graphs for Open-set Tasks

2026-05-29 · Marcel Bartholomeus Prasetyo, Shrutika Vishal Thengane, A Manicka Praveen, Yi Loo 외 arxiv

Open-set task execution can significantly benefit from seamlessly switching between coarse and fine scene representations depending on the context and the evolving information as the robot explores the environment. For e…

Scene Graph Generation

Learning Context-Conditioned Predicate Semantics via Prototype Feedback

2026-05-28 · NamGyu Jung, Chang Choi arxiv

In scene graph generation, a central challenge is modeling polysemous predicates whose meanings shift across contexts. Prior approaches address this issue by decomposing predicates into multiple static prototypes or retr…

Scene Graph Generation

RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses

2026-05-20 · Minh Anh Nguyen, Quang Huy Tran, Bao Ngoc Le, Tuan Kiet Pham 외 arxiv

Open-vocabulary 3D scene graph generation seeks to describe object instances and their relations with flexible natural-language predicates. The central difficulty is not only vocabulary expansion, but supervision reliabi…

Scene Graph Generation

Fixed External Cameras as Common Prior Maps for Active 3D Scene Graph Generation

2026-05-18 · Giorgia Modi, Davide Buoso, Giuseppe Averta, Daniele De Martini arxiv

Commonly available prior information, such as BIM models, floor plans, and remote sensing images, can provide valuable geometric and semantic context for autonomous robotic systems. In this paper, we treat observations f…

Scene Graph Generation3D Reconstruction

RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots

2026-05-18 · Giorgia Modi, Davide Buoso, Giuseppe Averta, Daniele De Martini arxiv

Current approaches to 3D scene graph generation rely on dedicated depth sensors, such as LiDAR or RGB-D cameras, for metric 3D reconstruction. This limits deployment to specialized robotic platforms and excludes settings…

Scene Graph Generation3D Reconstruction

Dependency-Aware Discrete Diffusion for Scene Graph Generation

2026-05-09 · Rajalaxmi Rajagopalan, Romit Roy Choudhury arxiv

Scene graphs (SGs) represent objects and their relationships as structured graphs, enabling applications in image generation, robotics, and 3D understanding. Recent work suggests that conditioning image generation on sce…

Scene Graph GenerationImage Generation

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding

2026-05-08 · Ke Ma, Jiaqi Tang, Bin Guo, Xueting Han 외 arxiv

Proactive streaming video understanding requires Video-LLMs to decide when to respond as a video unfolds, a task where existing methods often fall short due to their implicit, query-agnostic modeling of visual evidence. …

Scene Graph Generation
1–20 / 365 다음 →