paper-with-me

Papers

AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control

2026-03-15 · Peng Xu, Zhengnan Deng, Jiayan Deng, Zonghua Gu, Shaohua Wan arxiv

Vision-Language Navigation (VLN) for Unmanned Aerial Vehicles (UAVs) demands complex visual interpretation and continuous control in dynamic 3D environments. Existing hierarchical approaches rely on dense oracle guidance or auxiliary object detectors, creating semantic gaps and limiting genuine autonomy. We propose AerialVLA, a minimalist end-to-end Vision-Language-Action framework mapping raw visual observations and fuzzy linguistic instructions directly to continuous physical control signals. First, we introduce a streamlined dual-view perception strategy that reduces visual redundancy while preserving essential cues for forward navigation and precise grounding, which additionally facilitates future simulation-to-reality transfer. To reclaim genuine autonomy, we deploy a fuzzy directional prompting mechanism derived solely from onboard sensors, completely eliminating the dependency on dense oracle guidance. Ultimately, we formulate a unified control space that integrates continuous 3-Degree-of-Freedom (3-DoF) kinematic commands with an intrinsic landing signal, freeing the agent from external object detectors for precision landing. Extensive experiments on the TravelUAV benchmark demonstrate that AerialVLA achieves state-of-the-art performance in seen environments. Furthermore, it exhibits superior generalization in unseen scenarios by achieving nearly three times the success rate of leading baselines, validating that a minimalist, autonomy-centric paradigm captures more robust visual-motor representations than complex modular systems.

📄 PDF Abstract BibTeX arXiv:2603.14363

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language NavigationContinuous Control

Similar Papers 제목 키워드 기반

RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation

2026-08-10 · Boxiong Wang, Hui Kang, Geng Sun, Jiahui Li 외 arxiv

Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV…

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

2026-07-25 · Yihao Wu, Chenyi Xu, Liqi Yan, Chenhuan Cai 외 arxiv

Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based …

Vision-Language Navigation

Does Peer Observation Help? Vision-Sharing Collaboration for Vision-Language Navigation

2026-03-21 · Qunchao Jin, Yiliao Song, Qi Wu arxiv

Vision-Language Navigation (VLN) systems are fundamentally constrained by partial observability, as an agent can only accumulate knowledge from locations it has personally visited. As multiple robots increasingly coexist…

Vision-Language Navigation

Minimalist Vision with Freeform Pixels

2024-12-30 · Jeremy Klotz, Shree K. Nayar

A minimalist vision system uses the smallest number of pixels needed to solve a vision task. While traditional cameras use a large grid of square pixels, a minimalist camera uses freeform pixels that can take on arbitrar…

Vision-Dialog Navigation by Exploring Cross-modal Memory

2020-03-15 · CVPR 2020 6 · Yi Zhu, Fengda Zhu, Zhaohuan Zhan, Bingqian Lin 외

Vision-dialog navigation posed as a new holy-grail task in vision-language disciplinary targets at learning an agent endowed with the capability of constant conversation for help with natural language and navigating acco…

Decision Making