paper-with-me

홈 › Papers

ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation

2026-03-09 · Haoyu Tong, Xiangyu Dong, Xiaoguang Ma, Haoran Zhao, Yaoming Zhou, Chenghao Lin arxiv

Existing aerial Vision-Language Navigation (VLN) methods predominantly adopt a detection-and-planning pipeline, which converts open-vocabulary detections into discrete textual scene graphs. These approaches are plagued by inadequate spatial reasoning capabilities and inherent linguistic ambiguities. To address these bottlenecks, we propose a Visual-Spatial Reasoning (ViSA) enhanced framework for aerial VLN. Specifically, a triple-phase collaborative architecture is designed to leverage structured visual prompting, enabling Vision-Language Models (VLMs) to perform direct reasoning on image planes without the need for additional training or complex intermediate representations. Comprehensive evaluations on the CityNav benchmark demonstrate that the ViSA-enhanced VLN achieves a 70.3\% improvement in success rate compared to the fully trained state-of-the-art (SOTA) method, elucidating its great potential as a backbone for aerial VLN systems.

📄 PDF Abstract BibTeX arXiv:2603.08007

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language NavigationSpatial Reasoning

Similar Papers 제목 키워드 기반

History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation

2025-12-16 · Xichen Ding, Jianzhe Gao, Cong Pan, Wenguan Wang 외 arxiv

Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both …

AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations

2025-04-10 · Junli Liu, Qizhi Chen, Zhigang Wang, Yiwen Tang 외

Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual grounding from aerial views. Compared to traditi…

Spatial ReasoningVisual Grounding

View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification

2026-05-18 · Quan Zhang, Zeqiang Cai, Peiming Zhao, Jingze Wu 외 arxiv

Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existing methods typically follow a view-invariant paradigm, aligning shar…

Person Re-Identification

Local semantic enhanced convnet for aerial scene recognition

2021-07-08 · IEEE Transactions on Image Processing 2021 7 · Qi Bi, Kun Qin, Han Zhang, Gui-Song Xia

Aerial scene recognition is challenging due to the complicated object distribution and spatial arrangement in a large-scale aerial image. Recent studies attempt to explore the local semantic representation capability of …

Aerial Scene ClassificationImage ClassificationScene ClassificationScene Recognition

From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Pedagogical Visualization

2025-05-22 · Haonian Ji, Shi Qiu, Siyang Xin, Siwei Han 외

While foundation models (FMs), such as diffusion models and large vision-language models (LVLMs), have been widely applied in educational contexts, their ability to generate pedagogically effective visual explanations re…

Visual Reasoning