paper-with-me

Papers

IVLMap: Instance-Aware Visual Language Grounding for Consumer Robot Navigation

2024-03-28 · Jiacui Huang, Hongtao Zhang, Mingbo Zhao, Zhou Wu

Vision-and-Language Navigation (VLN) is a challenging task that requires a robot to navigate in photo-realistic environments with human natural language promptings. Recent studies aim to handle this task by constructing the semantic spatial map representation of the environment, and then leveraging the strong ability of reasoning in large language models for generalizing code for guiding the robot navigation. However, these methods face limitations in instance-level and attribute-level navigation tasks as they cannot distinguish different instances of the same object. To address this challenge, we propose a new method, namely, Instance-aware Visual Language Map (IVLMap), to empower the robot with instance-level and attribute-level semantic mapping, where it is autonomously constructed by fusing the RGBD video data collected from the robot agent with special-designed natural language map indexing in the bird's-in-eye view. Such indexing is instance-level and attribute-level. In particular, when integrated with a large language model, IVLMap demonstrates the capability to i) transform natural language into navigation targets with instance and attribute information, enabling precise localization, and ii) accomplish zero-shot end-to-end navigation tasks based on natural language commands. Extensive navigation experiments are conducted. Simulation results illustrate that our method can achieve an average improvement of 14.4\% in navigation accuracy. Code and demo are released at https://ivlmap.github.io/.

📄 PDF Abstract BibTeX arXiv:2403.19336

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeLanguage ModellingLarge Language ModelNavigateRobot NavigationVision and Language Navigation

Similar Papers 제목 키워드 기반

Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding

2024-11-05 · Sombit Dey, Ozan Unal, Christos Sakaridis, Luc van Gool

3D visual grounding consists of identifying the instance in a 3D scene which is referred by an accompanying language description. While several architectures have been proposed within the commonly employed grounding-by-s…

3D visual groundingVisual Grounding

Improving Generalized Visual Grounding with Instance-aware Joint Learning

2025-09-17 · Ming Dai, Wenxuan Cheng, Jiang-Jiang Liu, Lingfeng Yang 외 arxiv

Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target sce…

Generalized Referring Expression ComprehensionSemantic SegmentationVisual Grounding

GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

2026-08-19 · Chaowei Wang, Yan Di, Jingjun Sun, Baozhe Liu 외 arxiv

Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a natural language description. However, the abundance of small, densely distribut…

Spatial ReasoningVisual Grounding

3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment

2025-05-03 · Xiaoqi Li, Jiaming Liu, Nuowei Han, Liang Heng 외

The 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two pr…

SentenceVisual Grounding

InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual Referring

2021-03-01 · ICCV 2021 10 · Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang 외

Compared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer, to achieve a superior 3…

3D visual groundingAttributeObject LocalizationPanoptic Segmentation+1