paper-with-me

Papers

Reflex-Based Open-Vocabulary Navigation without Prior Knowledge Using Omnidirectional Camera and Multiple Vision-Language Models

2024-08-21 · Kento Kawaharazuka, Yoshiki Obinata, Naoaki Kanazawa, Naoto Tsukamoto, Kei Okada, Masayuki Inaba

Various robot navigation methods have been developed, but they are mainly based on Simultaneous Localization and Mapping (SLAM), reinforcement learning, etc., which require prior map construction or learning. In this study, we consider the simplest method that does not require any map construction or learning, and execute open-vocabulary navigation of robots without any prior knowledge to do this. We applied an omnidirectional camera and pre-trained vision-language models to the robot. The omnidirectional camera provides a uniform view of the surroundings, thus eliminating the need for complicated exploratory behaviors including trajectory generation. By applying multiple pre-trained vision-language models to this omnidirectional image and incorporating reflective behaviors, we show that navigation becomes simple and does not require any prior setup. Interesting properties and limitations of our method are discussed based on experiments with the mobile robot Fetch.

📄 PDF Abstract BibTeX arXiv:2408.11380

Code (0)

등록된 구현이 없습니다.

Tasks

Robot NavigationSimultaneous Localization and Mapping

Similar Papers 제목 키워드 기반

One Map to Find Them All: Real-time Open-Vocabulary Mapping for Zero-shot Multi-Object Navigation

2024-09-18 · Finn Lukas Busch, Timon Homberger, Jesús Ortega-Peimbert, Quantao Yang 외

The capability to efficiently search for objects in complex environments is fundamental for many real-world robot applications. Recent advances in open-vocabulary vision models have resulted in semantically-informed obje…

AllObject

HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation

2024-09-22 · Naoki Yokoyama, Ram Ramrakhya, Abhishek Das, Dhruv Batra 외

We present the Habitat-Matterport 3D Open Vocabulary Object Goal Navigation dataset (HM3D-OVON), a large-scale benchmark that broadens the scope and semantic range of prior Object Goal Navigation (ObjectNav) benchmarks. …

NavigateVisual Navigation

SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes

2025-05-24 · Dicong Qiu, Jiadi You, Zeying Gong, Ronghe Qiu 외

We present the Semantics-aware Dataset and Benchmark Generation Pipeline for Open-vocabulary Object Navigation in Dynamic Scenes (SD-OVON). It utilizes pretraining multimodal foundation models to generate infinite unique…

Object

DualMap: Online Open-Vocabulary Semantic Mapping for Natural Language Navigation in Dynamic Changing Scenes

2025-06-02 · Jiajun Jiang, Yiming Zhu, Zirui Wu, Jie Song

We introduce DualMap, an online open-vocabulary mapping system that enables robots to understand and navigate dynamically changing environments through natural language queries. Designed for efficient semantic mapping an…

Natural Language QueriesNavigateRobot Navigation

WildOS: Open-Vocabulary Object Search in the Wild

2026-02-22 · Hardik Shah, Erica Tevere, Deegan Atha, Marcel Kaufmann 외 arxiv

Autonomous navigation in complex, unstructured outdoor environments requires robots to operate over long ranges without prior maps and limited depth sensing. In such settings, relying solely on geometric frontiers for ex…

Visual Reasoning