paper-with-me

Papers

Hierarchical Language Models for Semantic Navigation and Manipulation in an Aerial-Ground Robotic System

2025-06-05 · Haokun Liu, Zhaoqi Ma, Yunong Li, Junichiro Sugihara, Yicheng Chen, Jinjie Li, Moju Zhao

Heterogeneous multi-robot systems show great potential in complex tasks requiring hybrid cooperation. However, traditional approaches relying on static models often struggle with task diversity and dynamic environments. This highlights the need for generalizable intelligence that can bridge high-level reasoning with low-level execution across heterogeneous agents. To address this, we propose a hierarchical framework integrating a prompted Large Language Model (LLM) and a GridMask-enhanced fine-tuned Vision Language Model (VLM). The LLM decomposes tasks and constructs a global semantic map, while the VLM extracts task-specified semantic labels and 2D spatial information from aerial images to support local planning. Within this framework, the aerial robot follows an optimized global semantic path and continuously provides bird-view images, guiding the ground robot's local semantic navigation and manipulation, including target-absent scenarios where implicit alignment is maintained. Experiments on real-world cube or object arrangement tasks demonstrate the framework's adaptability and robustness in dynamic environments. To the best of our knowledge, this is the first demonstration of an aerial-ground heterogeneous system integrating VLM-based perception with LLM-driven task reasoning and motion planning.

📄 PDF Abstract BibTeX arXiv:2506.05020

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelMotion Planning

Similar Papers 제목 키워드 기반

CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory

2025-05-08 · Weichen Zhang, Chen Gao, Shiquan Yu, Ruiying Peng 외

Aerial vision-and-language navigation (VLN), requiring drones to interpret natural language instructions and navigate complex urban environments, emerges as a critical embodied AI challenge that bridges human-robot inter…

Large Language ModelNavigateSpatial ReasoningVision and Language Navigation

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

2026-04-09 · Xingyu Xia, Lekai Zhou, Yujie Tang, Xiaozhou Zhu 외 arxiv

Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonomously navigate complex three-dimensional environments by grounding l…

Vision-Language Navigation

HANDO: Hierarchical Autonomous Navigation and Dexterous Omni-loco-manipulation

2025-10-10 · Jingyuan Sun, Chaoran Wang, Mingyu Zhang, Cui Miao 외 arxiv

Seamless loco-manipulation in unstructured environments requires robots to leverage autonomous exploration alongside whole-body control for physical interaction. In this work, we introduce HANDO (Hierarchical Autonomous …

APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation

2026-01-31 · Daoxuan Zhang, Ping Chen, Xiaobo Xia, Xiu Su 외 arxiv

Aerial Object Goal Navigation, a challenging frontier in Embodied AI, requires an Unmanned Aerial Vehicle (UAV) agent to autonomously explore, reason, and identify a specific target using only visual perception and langu…

Reinforcement Learning

DroneVLA: VLA based Aerial Manipulation

2026-01-20 · Fawad Mehboob, Monijesu James, Amir Habel, Jeffrin Sam 외 arxiv

As aerial platforms evolve from passive observers to active manipulators, the challenge shifts toward designing intuitive interfaces that allow non-expert users to command these systems naturally. This work introduces a …

Pose Estimation