paper-with-me

Papers

Towards Versatile Embodied Navigation

2022-10-30 · Hanqing Wang, Wei Liang, Luc van Gool, Wenguan Wang

With the emergence of varied visual navigation tasks (e.g, image-/object-/audio-goal and vision-language navigation) that specify the target in different ways, the community has made appealing advances in training specialized agents capable of handling individual navigation tasks well. Given plenty of embodied navigation tasks and task-specific solutions, we address a more fundamental question: can we learn a single powerful agent that masters not one but multiple navigation tasks concurrently? First, we propose VXN, a large-scale 3D dataset that instantiates four classic navigation tasks in standardized, continuous, and audiovisual-rich environments. Second, we propose Vienna, a versatile embodied navigation agent that simultaneously learns to perform the four navigation tasks with one model. Building upon a full-attentive architecture, Vienna formulates various navigation tasks as a unified, parse-and-query procedure: the target description, augmented with four task embeddings, is comprehensively interpreted into a set of diversified goal vectors, which are refined as the navigation progresses, and used as queries to retrieve supportive context from episodic history for decision making. This enables the reuse of knowledge across navigation tasks with varying input domains/modalities. We empirically demonstrate that, compared with learning each visual navigation task individually, our multitask agent achieves comparable or even better performance with reduced complexity.

📄 PDF Abstract BibTeX arXiv:2210.16822

Code (1)

hanqingwangai/vxn 공식 구현 pytorch

Tasks

Decision MakingVision-Language NavigationVisual Navigation

Similar Papers 제목 키워드 기반

VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation

2025-12-22 · Sihao Lin, Zerui Li, Xunyi Zhao, Gengze Zhou 외 arxiv

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomings limit the insight that the benchmarks…

Vision-Language Navigation

SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation

2026-03-10 · Jiahang Liu, Tianyu Xu, Jiawei Chen, Lu Yue 외 arxiv

Recent embodied navigation approaches leveraging Vision-Language Models (VLMs) demonstrate strong generalization in versatile Vision-Language Navigation (VLN). However, reliable path planning in complex environments rema…

Vision-Language Navigation

Communicative Learning with Natural Gestures for Embodied Navigation Agents with Human-in-the-Scene

2021-08-05 · Qi Wu, Cheng-Ju Wu, Yixin Zhu, Jungseock Joo

Human-robot collaboration is an essential research topic in artificial intelligence (AI), enabling researchers to devise cognitive AI systems and affords an intuitive means for users to interact with the robot. Of note, …

ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation

2026-02-12 · Zedong Chu, Shichao Xie, Xiaolong Wu, Yanfen Shen 외 arxiv

Embodied navigation has long been fragmented by task-specific architectures. We introduce ABot-N0, a unified Vision-Language-Action (VLA) foundation model that achieves a ``Grand Unification'' across 5 core tasks: Point-…

From reactive to cognitive: brain-inspired spatial intelligence for embodied agents

2025-08-24 · Shouwei Ruan, Liyuan Wang, Caixin Kang, Qihui Zhu 외 arxiv

Spatial cognition enables adaptive goal-directed behavior by constructing internal models of space. Robust biological systems consolidate spatial knowledge into three interconnected forms: \textit{landmarks} for salient …

Zero-shot Generalization