paper-with-me

Papers

Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

2025-06-12 · Yuhang Zhang, Haosheng Yu, Jiaping Xiao, Mir Feroskhan

Vision-and-language navigation (VLN) is a long-standing challenge in autonomous robotics, aiming to empower agents with the ability to follow human instructions while navigating complex environments. Two key bottlenecks remain in this field: generalization to out-of-distribution environments and reliance on fixed discrete action spaces. To address these challenges, we propose Vision-Language Fly (VLFly), a framework tailored for Unmanned Aerial Vehicles (UAVs) to execute language-guided flight. Without the requirement for localization or active ranging sensors, VLFly outputs continuous velocity commands purely from egocentric observations captured by an onboard monocular camera. The VLFly integrates three modules: an instruction encoder based on a large language model (LLM) that reformulates high-level language into structured prompts, a goal retriever powered by a vision-language model (VLM) that matches these prompts to goal images via vision-language similarity, and a waypoint planner that generates executable trajectories for real-time UAV control. VLFly is evaluated across diverse simulation environments without additional fine-tuning and consistently outperforms all baselines. Moreover, real-world VLN tasks in indoor and outdoor environments under direct and indirect instructions demonstrate that VLFly achieves robust open-vocabulary goal understanding and generalized navigation capabilities, even in the presence of abstract language input.

📄 PDF Abstract BibTeX arXiv:2506.10756

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelVision and Language NavigationVision-Language Navigation

Similar Papers 제목 키워드 기반

IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments

2025-12-22 · Xu Liu, Yu Liu, Hanshuo Qiu, Yang Qirong 외 arxiv

Vision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robo…

Vision-Language NavigationMultimodal ReasoningData Augmentation

OpenVLN: Open-world Aerial Vision-Language Navigation

2025-11-09 · Peican Lin, Gan Sun, Chenxi Liu, Fazeng Li 외 arxiv

Vision-language models (VLMs) have been widely-applied in ground-based vision-language navigation (VLN). However, the vast complexity of outdoor aerial environments compounds data acquisition challenges and imposes long-…

Vision-Language NavigationReinforcement LearningTrajectory Planning

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

2026-04-09 · Xingyu Xia, Lekai Zhou, Yujie Tang, Xiaozhou Zhu 외 arxiv

Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonomously navigate complex three-dimensional environments by grounding l…

Vision-Language Navigation

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap

2026-04-15 · Hanxuan Chen, Jie Zheng, Siqi Yang, Tianle Zeng 외 arxiv

Vision-and-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) represents a pivotal challenge in embodied artificial intelligence, focused on enabling UAVs to interpret high-level human commands and execute long-h…

SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments

2025-07-09 · Tianshun Li, Tianyi Huai, Zhen Li, Yichun Gao 외 arxiv

Unmanned Aerial Vehicles (UAVs) have emerged as versatile tools across various sectors, driven by their mobility and adaptability. This paper introduces SkyVLN, a novel framework integrating vision-and-language navigatio…