paper-with-me

Papers

Semantic-Aware Guided Drone Exploration for Language-Conditioned 3D Indoor Mapping

2026-05-22 · Nitin Vegesna, Avideh Zakhor arxiv

We present Semantic-Aware Guided Exploration, SAGE, a system for open-vocabulary exploration in unknown 3D indoor environments that preserves coverage-oriented behavior while allowing semantic cues to reprioritize frontier selection. Building on the FALCON volumetric explorer, SAGE integrates Contrastive Language-Image Pre-training (CLIP) via four key components: object-centric embedding storage, a temporal cache that projects recent observations onto the free-unknown boundary, object frontiers for high-similarity detections, and a unified semantic-geometric planning cost. This cost function bounds semantic reweighting influence, ensuring frontiers are prioritized without sacrificing total coverage. In Matterport3D-based simulations, SAGE outperforms FALCON and a semantic-only ablation in object discovery across map-query pairs. Compared to Finding Things in the Unknown (FTU), SAGE completes exploration 9.0 to 25.9 times faster across the nine shared map-query pairs, achieving a mean speedup of 13.7. Furthermore, SAGE achieves substantially higher volumetric throughput than FTU. Finally, we deploy SAGE in five real-world flights in two environments on a Modal AI Starling 2 quadrotor with onboard sensing and planning, and offboard CLIP inference. Comparing SAGE and FALCON, we find that while FALCON results in faster exploration and shorter mapping trajectories, SAGE outperforms FALCON in terms of object discovery.

📄 PDF Abstract BibTeX arXiv:2605.23160

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

2026-08-27 · Jiahao Wen, Hang Yu, Zhedong Zheng arxiv

Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task as direct matching between an open-ended…

VLN-Pilot: Large Vision-Language Model as an Autonomous Indoor Drone Operator

2026-02-05 · Bessie Dominguez-Dager, Sergio Suescun-Ferrandiz, Felix Escalona, Francisco Gomez-Donoso 외 arxiv

This paper introduces VLN-Pilot, a novel framework in which a large Vision-and-Language Model (VLLM) assumes the role of a human pilot for indoor drone navigation. By leveraging the multimodal reasoning abilities of VLLM…

Multimodal ReasoningDrone navigation

HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones

2025-08-29 · Hao Ruan, Jinliang Lin, Yingxin Lai, Zhiming Luo 외 arxiv

Natural Language-Guided Drones (NLGD) provide a novel paradigm for tasks such as target matching and navigation. However, the wide field of view and complex compositional semantics in drone scenarios pose challenges for …

Zero-shot GeneralizationContrastive LearningImage-text matchingImage Retrieval

Guided Navigation in Knowledge-Dense Environments: Structured Semantic Exploration with Guidance Graphs

2025-08-06 · Dehao Tao, Guangjie Liu, Weizheng, Yongfeng Huang 외 arxiv

While Large Language Models (LLMs) exhibit strong linguistic capabilities, their reliance on static knowledge and opaque reasoning processes limits their performance in knowledge intensive tasks. Knowledge graphs (KGs) o…

Knowledge Graphs

Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4

2025-10-03 · Lingfeng Zhang, Erjia Xiao, Yuchen Zhang, Haoxiang Fu 외 arxiv

Cross-modal drone navigation remains a challenging task in robotics, requiring efficient retrieval of relevant images from large-scale databases based on natural language descriptions. The RoboSense 2025 Track 4 challeng…

Drone navigationImage Retrieval