paper-with-me

홈 › Papers

DA-Nav: Direction-Aware City-Scale Vision-Language Navigation

2026-07-13 · Ye Yuan, Kehan Chen, Xinqiang Yu, Wentao Xu, Heng Wang, Libo Huang, Chuanguang Yang, Yan Huang, Jiawei He, Zhulin An arxiv

City-scale outdoor navigation is currently hindered by the heavy reliance on dense maps or costly navigation supervision. In this work, we introduce a novel paradigm for leveraging directional instructions from commercial navigation tools (e.g., Google Maps). To bridge the gap between commercial instructions and executable navigation actions, while mitigating long-horizon error accumulation through robust trajectory recovery, we propose DA-Nav, a Direction-Aware vision-language Navigation framework that reformulates navigation as a discrete spatial grounding problem on the egocentric 2D image plane. To achieve trajectory recovery, DA-Nav employs a Chain-of-Thought (CoT) reasoning process encompassing deviation assessment, action prediction, and target grid selection. We further introduce ReDA, a dataset that provides direction-aware instructions and recovery trajectories to enhance spatial grounding and support CoT recovery reasoning. Extensive experiments in CARLA demonstrate that DA-Nav achieves a high success rate of 56.16% in unseen urban environments, outperforming existing State-of-The-Art (SoTA) methods while maintaining a substantially stronger recovery capability. Furthermore, without fine-tuning, DA-Nav seamlessly adapts to both quadruped and humanoid robots, enabling stable kilometer-scale closed-loop outdoor navigation in complex real world environments.

📄 PDF Abstract BibTeX arXiv:2607.11638

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language Navigation

Similar Papers 제목 키워드 기반

StyleCity: Large-Scale 3D Urban Scenes Stylization

2024-04-16 · Yingshu Chen, Huajian Huang, Tuan-Anh Vu, Ka Chun Shum 외

Creating large-scale virtual urban scenes with variant styles is inherently challenging. To facilitate prototypes of virtual production and bypass the need for complex materials and lighting setups, we introduce the firs…

LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

2026-04-19 · Yuwei Ning, Ganlong Zhao, Yipeng Qin, Si Liu 외 arxiv

Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex urban environments. While recent advances have achieved progress thr…

Computational EfficiencySpatial Reasoning

No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models

2025-10-04 · Min Woo Sun, Alejandro Lozano, Javier Gamazo Tejero, Vishwesh Nath 외 arxiv

Embedding vision-language models (VLMs) are typically pretrained with short text windows (<77 tokens), which forces the truncation of long-format captions. Yet, the distribution of biomedical captions from large-scale op…

Language-guided Scale-aware MedSegmentor for Lesion Segmentation in Medical Imaging

2024-08-30 · Shuyi Ouyang, Jinyang Zhang, Xiangye Lin, Xilai Wang 외

In clinical practice, segmenting specific lesions based on the needs of physicians can significantly enhance diagnostic accuracy and treatment efficiency. However, conventional lesion segmentation models lack the flexibi…

DiagnosticImage SegmentationLanguage ModelingLanguage Modelling+5

3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding

2026-03-24 · Yiping Chen, Jinpeng Li, Wenyu Ke, Yang Luo 외 arxiv

While multi-modality large language models excel in object-centric or indoor scenarios, scaling them to 3D city-scale environments remains a formidable challenge. To bridge this gap, we propose 3DCity-LLM, a unified fram…

Spatial Reasoning