paper-with-me

Papers

Intelligent Spatial Perception by Building Hierarchical 3D Scene Graphs for Indoor Scenarios with the Help of LLMs

2025-03-19 · Yao Cheng, Zhe Han, Fengyang Jiang, Huaizhen Wang, Fengyu Zhou, Qingshan Yin, Lei Wei

This paper addresses the high demand in advanced intelligent robot navigation for a more holistic understanding of spatial environments, by introducing a novel system that harnesses the capabilities of Large Language Models (LLMs) to construct hierarchical 3D Scene Graphs (3DSGs) for indoor scenarios. The proposed framework constructs 3DSGs consisting of a fundamental layer with rich metric-semantic information, an object layer featuring precise point-cloud representation of object nodes as well as visual descriptors, and higher layers of room, floor, and building nodes. Thanks to the innovative application of LLMs, not only object nodes but also nodes of higher layers, e.g., room nodes, are annotated in an intelligent and accurate manner. A polling mechanism for room classification using LLMs is proposed to enhance the accuracy and reliability of the room node annotation. Thorough numerical experiments demonstrate the system's ability to integrate semantic descriptions with geometric data, creating an accurate and comprehensive representation of the environment instrumental for context-aware navigation and task planning.

📄 PDF Abstract BibTeX arXiv:2503.15091

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectRobot NavigationTask Planning

Similar Papers 제목 키워드 기반

HSPFormer: Hierarchical Spatial Perception Transformer for Semantic Segmentation

2025-01-16 · IEEE Transactions on Intelligent Transportation Systems 2025 1 · Siyu Chen, Ting Han, Changshe Zhang, Jinhe Su 외

Semantic perception in driving scenarios plays a crucial role in intelligent transportation systems. However, existing Transformer-based semantic segmentation methods often do not fully exploit their potential in underst…

Depth EstimationMonocular Depth EstimationSemantic SegmentationSpatial Reasoning

ROOT: VLM based System for Indoor Scene Understanding and Beyond

2024-11-24 · Yonghui Wang, Shi-Yong Chen, Zhenxing Zhou, Siyi Li 외

Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In this study, we introduce ROOT, a VLM-ba…

Scene GenerationScene Understanding

Reconfigurable Intelligent Surface Aided Wireless Sensing for Scene Depth Estimation

2022-11-15 · Abdelrahman Taha, Hao Luo, Ahmed Alkhateeb

Current scene depth estimation approaches mainly rely on optical sensing, which carries privacy concerns and suffers from estimation ambiguity for distant, shiny, and transparent surfaces/objects. Reconfigurable intellig…

Depth Estimation

CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping

2026-04-13 · Yiran Ling, Wenxuan Li, Siying Dong, Yize Zhang 외 arxiv

Robot grasping of desktop object is widely used in intelligent manufacturing, logistics, and agriculture.Although vision-language models (VLMs) show strong potential for robotic manipulation, their deployment in low-leve…

Logical Reasoning

3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans

2020-02-15 · Antoni Rosinol, Arjun Gupta, Marcus Abate, Jingnan Shi 외

We present a unified representation for actionable spatial perception: 3D Dynamic Scene Graphs. Scene graphs are directed graphs where nodes represent entities in the scene (e.g. objects, walls, rooms), and edges represe…

3D ReconstructionHuman DetectionRobot Task PlanningSimultaneous Localization and Mapping+1