paper-with-me

홈 › Papers

Context-Aware Entity Grounding with Open-Vocabulary 3D Scene Graphs

2023-09-27 · Haonan Chang, Kowndinya Boyalakuntla, Shiyang Lu, Siwei Cai, Eric Jing, Shreesh Keskar, Shijie Geng, Adeeb Abbas, Lifeng Zhou, Kostas Bekris, Abdeslam Boularias

We present an Open-Vocabulary 3D Scene Graph (OVSG), a formal framework for grounding a variety of entities, such as object instances, agents, and regions, with free-form text-based queries. Unlike conventional semantic-based object localization approaches, our system facilitates context-aware entity localization, allowing for queries such as `pick up a cup on a kitchen table" or `navigate to a sofa on which someone is sitting". In contrast to existing research on 3D scene graphs, OVSG supports free-form text input and open-vocabulary querying. Through a series of comparative experiments using the ScanNet dataset and a self-collected dataset, we demonstrate that our proposed approach significantly surpasses the performance of previous semantic-based localization techniques. Moreover, we highlight the practical application of OVSG in real-world robot navigation and manipulation experiments.

📄 PDF Abstract BibTeX arXiv:2309.15940

Code (1)

changhaonan/ovsg 공식 구현 pytorch

Tasks

FormNavigateObjectObject LocalizationRobot Navigation

Similar Papers 제목 키워드 기반

A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection

2026-08-28 · Zhoupeng Guo, Xinjie Yao, Yunqi Zhu, Zhihe Fan 외 arxiv

Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelligence, grounding a language command into the same physical target bef…

Referring Expression

Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding

2025-09-08 · Jiangnan Xie, Xiaolong Zheng, Liang Zheng arxiv

Visual Grounding (VG) aims to utilize given natural language queries to locate specific target objects within images. While current transformer-based approaches demonstrate strong localization performance in standard sce…

Natural Language QueriesVisual Grounding

Part-Aware Open-Vocabulary 3D Affordance Grounding via Prototypical Semantic and Geometric Alignment

2026-03-18 · Dongqiang Gou, Xuming He arxiv

Grounding natural language questions to functionally relevant regions in 3D objects -- termed language-driven 3D affordance grounding -- is essential for embodied intelligence and human-AI interaction. Existing methods, …

Betrayed by Captions: Joint Caption Grounding and Generation for Open Vocabulary Instance Segmentation

2023-01-02 · ICCV 2023 1 · Jianzong Wu, Xiangtai Li, Henghui Ding, Xia Li 외

In this work, we focus on open vocabulary instance segmentation to expand a segmentation model to classify and segment instance-level novel categories. Previous approaches have relied on massive caption datasets and comp…

Caption GenerationInstance SegmentationPanoptic SegmentationSegmentation+1

Open-Vocabulary 3D Semantic Segmentation with Foundation Models

2024-01-01 · CVPR 2024 1 · Li Jiang, Shaoshuai Shi, Bernt Schiele

In dynamic 3D environments the ability to recognize a diverse range of objects without the constraints of predefined categories is indispensable for real-world applications. In response to this need we introduce OV3D…

3D Semantic SegmentationOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentation+2