paper-with-me

홈 › Papers

Can Vision Foundation Models Navigate? Zero-Shot Real-World Evaluation and Lessons Learned

2026-03-26 · Maeva Guerrier, Karthik Soma, Jana Pavlasek, Giovanni Beltrame arxiv

Visual Navigation Models (VNMs) promise generalizable, robot navigation by learning from large-scale visual demonstrations. Despite growing real-world deployment, existing evaluations rely almost exclusively on success rate, whether the robot reaches its goal, which conceals trajectory quality, collision behavior, and robustness to environmental change. We present a real-world evaluation of five state-of-the-art VNMs (GNM, ViNT, NoMaD, NaviBridger, and CrossFormer) across two robot platforms and five environments spanning indoor and outdoor settings. Beyond success rate, we combine path-based metrics with vision-based goal-recognition scores and assess robustness through controlled image perturbations (motion blur, sunflare). Our analysis uncovers three systematic limitations: (a) even architecturally sophisticated diffusion and transformer-based models exhibit frequent collisions, indicating limited geometric understanding; (b) models fail to discriminate between different locations that are perceptually similar, however some semantics differences are present, causing goal prediction errors in repetitive environments; and (c) performance degrades under distribution shift. We will publicly release our evaluation codebase and dataset to facilitate reproducible benchmarking of VNMs.

📄 PDF Abstract BibTeX arXiv:2603.25937

Code (0)

등록된 구현이 없습니다.

Tasks

Visual NavigationRobot Navigation

Similar Papers 제목 키워드 기반

OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision-Language Foundation Models

2024-02-16 · Yuxuan Kuang, Hai Lin, Meng Jiang

Object navigation (ObjectNav) requires an agent to navigate through unseen environments to find queried objects. Many previous methods attempted to solve this task by relying on supervised or reinforcement learning, wher…

Common Sense ReasoningNavigate

$A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

2023-08-15 · Peihao Chen, Xinyu Sun, Hongyan Zhi, Runhao Zeng 외

We study the task of zero-shot vision-and-language navigation (ZS-VLN), a practical yet challenging problem in which an agent learns to navigate following a path described by language instructions without requiring any p…

NavigateRobot NavigationVision and Language Navigation

VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation

2023-12-06 · Naoki Yokoyama, Sehoon Ha, Dhruv Batra, Jiuguang Wang 외

Understanding how humans leverage semantic knowledge to navigate unfamiliar environments and decide where to explore next is pivotal for developing robots capable of human-like search behaviors. We introduce a zero-shot …

Language ModellingNavigate

ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts

2025-07-06 · Sangbum Choi, Kyeongryeol Go, Taewoong Jang arxiv

Foundation models have revolutionized AI, yet they struggle with zero-shot deployment in real-world industrial settings due to a lack of high-quality, domain-specific datasets. To bridge this gap, Superb AI introduces ZE…

Few-Shot Object Detection

FoundationStereo: Zero-Shot Stereo Matching

2025-01-17 · CVPR 2025 1 · Bowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz 외

Tremendous progress has been made in deep stereo matching to excel on benchmark datasets through per-domain fine-tuning. However, achieving strong zero-shot generalization - a hallmark of foundation models in other compu…

Depth EstimationDiversityStereo Depth EstimationStereo Matching+1