paper-with-me

홈 › Papers

Open-Set 3D Semantic Instance Maps for Vision Language Navigation -- O3D-SIM

2024-04-27 · Laksh Nanwani, Kumaraditya Gupta, Aditya Mathur, Swayam Agrawal, A. H. Abdul Hafez, K. Madhava Krishna

Humans excel at forming mental maps of their surroundings, equipping them to understand object relationships and navigate based on language queries. Our previous work SI Maps [1] showed that having instance-level information and the semantic understanding of an environment helps significantly improve performance for language-guided tasks. We extend this instance-level approach to 3D while increasing the pipeline's robustness and improving quantitative and qualitative results. Our method leverages foundational models for object recognition, image segmentation, and feature extraction. We propose a representation that results in a 3D point cloud map with instance-level embeddings, which bring in the semantic understanding that natural language commands can query. Quantitatively, the work improves upon the success rate of language-guided tasks. At the same time, we qualitatively observe the ability to identify instances more clearly and leverage the foundational models and language and image-aligned embeddings to identify objects that, otherwise, a closed-set approach wouldn't be able to identify.

📄 PDF Abstract BibTeX arXiv:2404.17922

Code (1)

Smart-Wheelchair-RRC/o3d-sim 공식 구현 pytorch

Tasks

Image SegmentationNavigateObject RecognitionSemantic SegmentationVision-Language Navigation

Similar Papers 제목 키워드 기반

RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation

2025-05-21 · Naman Patel, Prashanth Krishnamurthy, Farshad Khorrami

Mapping and understanding complex 3D environments is fundamental to how autonomous systems perceive and interact with the physical world, requiring both precise geometric reconstruction and rich semantic comprehension. W…

GPUNatural Language QueriesObjectobject-detection+3

FM-Fusion: Instance-aware Semantic Mapping Boosted by Vision-Language Foundation Models

2024-02-07 · Chuhao Liu, Ke Wang, Jieqi Shi, Zhijian Qiao 외

Semantic mapping based on the supervised object detectors is sensitive to image distribution. In real-world environments, the object detection and segmentation performance can lead to a major drop, preventing the use of …

Instance SegmentationObjectobject-detectionObject Detection+2

FUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic Mapping by 3D Fusion of Voxel- and Instance-Level Layers

2026-05-05 · Timon Homberger, Finn Lukas Busch, Jesús Gerardo Ortega Peimbert, Quantao Yang 외 arxiv

Open-vocabulary semantic mapping enables robots to spatially ground previously unseen concepts without requiring predefined class sets. Current training-free methods commonly rely on multi-view fusion of semantic embeddi…

3D Semantic Segmentation

OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping

2025-08-03 · Danyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang 외 arxiv

Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic repre…

FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment

2025-04-11 · Sebastián Barbas Laina, Simon Boche, Sotiris Papatheodorou, Simon Schaefer 외

Geometrically accurate and semantically expressive map representations have proven invaluable to facilitate robust and safe mobile robot navigation and task planning. Nevertheless, real-time, open-vocabulary semantic und…

3D geometryNatural Language QueriesRobot NavigationScene Understanding+1