paper-with-me

Papers

Aerial Vision-and-Language Navigation with Grid-based View Selection and Map Construction

2025-03-14 · Ganlong Zhao, Guanbin Li, Jia Pan, Yizhou Yu

Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following human instruction. Compared to ground-based VLN, aerial VLN requires the agent to decide the next action in both horizontal and vertical directions based on the first-person view observations. Previous methods struggle to perform well due to the longer navigation path, more complicated 3D scenes, and the neglect of the interplay between vertical and horizontal actions. In this paper, we propose a novel grid-based view selection framework that formulates aerial VLN action prediction as a grid-based view selection task, incorporating vertical action prediction in a manner that accounts for the coupling with horizontal actions, thereby enabling effective altitude adjustments. We further introduce a grid-based bird's eye view map for aerial space to fuse the visual information in the navigation history, provide contextual scene information, and mitigate the impact of obstacles. Finally, a cross-modal transformer is adopted to explicitly align the long navigation history with the instruction. We demonstrate the superiority of our method in extensive experiments.

📄 PDF Abstract BibTeX arXiv:2503.11091

Code (0)

등록된 구현이 없습니다.

Tasks

NavigateVision and Language Navigation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Hierarchical Language Models for Semantic Navigation and Manipulation in an Aerial-Ground Robotic System

2025-06-05 · Haokun Liu, Zhaoqi Ma, Yunong Li, Junichiro Sugihara 외

Heterogeneous multi-robot systems show great potential in complex tasks requiring hybrid cooperation. However, traditional approaches relying on static models often struggle with task diversity and dynamic environments. …

Language ModelingLanguage ModellingLarge Language ModelMotion Planning

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

2026-04-09 · Xingyu Xia, Lekai Zhou, Yujie Tang, Xiaozhou Zhu 외 arxiv

Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonomously navigate complex three-dimensional environments by grounding l…

Vision-Language Navigation

History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation

2025-12-16 · Xichen Ding, Jianzhe Gao, Cong Pan, Wenguan Wang 외 arxiv

Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both …

OpenVLN: Open-world Aerial Vision-Language Navigation

2025-11-09 · Peican Lin, Gan Sun, Chenxi Liu, Fazeng Li 외 arxiv

Vision-language models (VLMs) have been widely-applied in ground-based vision-language navigation (VLN). However, the vast complexity of outdoor aerial environments compounds data acquisition challenges and imposes long-…

Vision-Language NavigationReinforcement LearningTrajectory Planning

FreqNav: Stage-Wise Frequency Routing for Object-Oriented Aerial Vision-Language Navigation

2026-08-02 · Yin Tang, Jiawei Ma, Jiahao Li, Hao Zhang 외 arxiv

Object-oriented aerial vision-and-language navigation (VLN) requires searching for a described target and landing on it precisely, under long-horizon and closed-loop control. Guided by a target-descriptive instruction du…

Vision-Language NavigationContinuous Control