paper-with-me

홈 › Papers

Less is More: Toward Zero-Shot Local Scene Graph Generation via Foundation Models

2023-10-02 · Shu Zhao, Huijuan Xu

Humans inherently recognize objects via selective visual perception, transform specific regions from the visual field into structured symbolic knowledge, and reason their relationships among regions based on the allocation of limited attention resources in line with humans' goals. While it is intuitive for humans, contemporary perception systems falter in extracting structural information due to the intricate cognitive abilities and commonsense knowledge required. To fill this gap, we present a new task called Local Scene Graph Generation. Distinct from the conventional scene graph generation task, which encompasses generating all objects and relationships in an image, our proposed task aims to abstract pertinent structural information with partial objects and their relationships for boosting downstream tasks that demand advanced comprehension and reasoning capabilities. Correspondingly, we introduce zEro-shot Local scEne GrAph geNeraTion (ELEGANT), a framework harnessing foundation models renowned for their powerful perception and commonsense reasoning, where collaboration and information communication among foundation models yield superior outcomes and realize zero-shot local scene graph generation without requiring labeled supervision. Furthermore, we propose a novel open-ended evaluation metric, Entity-level CLIPScorE (ECLIPSE), surpassing previous closed-set evaluation metrics by transcending their limited label space, offering a broader assessment. Experiment results show that our approach markedly outperforms baselines in the open-ended evaluation setting, and it also achieves a significant performance boost of up to 24.58% over prior methods in the close-set setting, demonstrating the effectiveness and powerful reasoning ability of our proposed framework.

📄 PDF Abstract BibTeX arXiv:2310.01356

Code (1)

bowen-upenn/Multi-Agent-VQA pytorch

Tasks

Graph GenerationScene Graph Generation

Similar Papers 제목 키워드 기반

PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory

2025-11-10 · Qunchao Jin, Yilin Wu, Changhao Chen arxiv

Zero-shot object navigation (ZSON) in unseen environments remains a challenging problem for household robots, requiring strong perceptual understanding and decision-making capabilities. While recent methods leverage metr…

Spatial ReasoningScene Parsing

BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation

2024-07-25 · Xiang Zhang, Bingxin Ke, Hayko Riemenschneider, Nando Metzger 외

By training over large-scale datasets, zero-shot monocular depth estimation (MDE) methods show robust performance in the wild but often suffer from insufficient detail. Although recent diffusion-based MDE approaches exhi…

Depth EstimationMonocular Depth Estimation

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

2025-08-28 · Jiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan 외 arxiv

3D Visual Grounding (3DVG) aims to localize objects in 3D scenes using natural language descriptions. Although supervised methods achieve higher accuracy in constrained settings, zero-shot 3DVG holds greater promise for …

3D Semantic SegmentationVisual Grounding

SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation

2026-01-11 · Jiwen Zhang, Zejun Li, Siyuan Wang, Xiangyu Shi 외 arxiv

Although learning-based vision-and-language navigation (VLN) agents can learn spatial knowledge implicitly from large-scale training data, zero-shot VLN agents lack this process, relying primarily on local observations f…

Object Localization

ConRF: Zero-shot Stylization of 3D Scenes with Conditioned Radiation Fields

2024-02-02 · Xingyu Miao, Yang Bai, Haoran Duan, Fan Wan 외

Most of the existing works on arbitrary 3D NeRF style transfer required retraining on each single style condition. This work aims to achieve zero-shot controlled stylization in 3D scenes utilizing text or visual input as…

NeRFStyle Transfer