paper-with-me

홈 › Papers

OpenFrontier: General Navigation with Visual-Language Grounded Frontiers

2026-03-05 · Esteban Padilla-Cerdio, Boyang Sun, Marc Pollefeys, Hermann Blum arxiv

Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Conventional navigation approaches often rely on dense 3D reconstruction and hand-crafted goal metrics, which limits their generalization across tasks and environments. Recent advances in vision-language navigation (VLN) and vision-language-action (VLA) models enable end-to-end policies conditioned on natural language, but typically require interactive training, large-scale data collection, or task-specific fine-tuning with a mobile agent. We formulate navigation as a sparse subgoal identification and reaching problem and observe that providing visual anchoring targets for high-level semantic priors enables highly efficient goal-conditioned navigation. Based on this insight, we select visual frontiers as semantic anchors and propose OpenFrontier, a navigation framework that requires no task-specific training or fine-tuning and seamlessly integrates diverse vision-language prior models. OpenFrontier enables efficient navigation with a lightweight system design, without dense 3D semantic mapping, task-specific policy training, or model fine-tuning. We evaluate OpenFrontier across multiple navigation benchmarks and demonstrate strong zero-shot performance, as well as effective real-world deployment on a mobile robot.

📄 PDF Abstract BibTeX arXiv:2603.05377

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language Navigation3D Reconstruction

Similar Papers 제목 키워드 기반

Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments

2017-11-20 · CVPR 2018 6 · Peter Anderson, Qi Wu, Damien Teney, Jake Bruce 외

A robot that can carry out a natural-language instruction has been a dream since before the Jetsons cartoon series imagined a life of leisure mediated by a fleet of attentive robot helpers. It is a dream that remains stu…

Reinforcement LearningTranslationVision and Language NavigationVisual Navigation+2

ABot-N1: Toward a General Visual Language Navigation Foundation Model

2026-07-11 · Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong 외 arxiv

Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks. Current approaches typically achieve this integration via monolit…

Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas

2022-03-25 · ACL 2022 5 · Raphael Schumann, Stefan Riezler

Vision and language navigation (VLN) is a challenging visually-grounded language understanding task. Given a natural language navigation instruction, a visual agent interacts with a graph-based environment equipped with …

DiversityVision and Language Navigation

\textsc{NaVIDA}: Vision-Language Navigation with Inverse Dynamics Augmentation

2026-01-26 · Weiye Zhu, Zekai Zhang, Xiangchen Wang, Hewei Pan 외 arxiv

Vision-and-Language Navigation (VLN) requires agents to interpret natural language instructions and act coherently in visually rich environments. However, most existing methods rely on reactive state-action mappings with…

Vision-Language Navigation

Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation

2024-03-26 · Abdelrhman Werby, Chenguang Huang, Martin Büchner, Abhinav Valada 외

Recent open-vocabulary robot mapping methods enrich dense geometric maps with pre-trained visual-language features. While these maps allow for the prediction of point-wise saliency maps when queried for a certain languag…

ObjectRobot Navigation