paper-with-me

홈 › Papers

Open-Vocabulary Object-Goal Navigation by Generalizing Semantic Mapping with Dense CLIP

2024-07-12 · Meng Wei, Chenyang Wan, Tai Wang, Yuqiang Yang, Wenzhe Cai, Yilun Chen, Hanqing Wang, Jiangmiao Pang, Xihui Liu arxiv

Object-oriented embodied navigation tasks require agents to locate specific objects, either defined by category or images, in unseen environments. While recent methods have made progress in extending closed-set models to open-vocabulary scenarios with foundation models, they typically rely on training-free large language models (LLMs) or finetuning with end-to-end reinforcement learning (RL). However, they face challenges in efficiency (e.g., the overhead and cost of LLM inference) and limited generalization from intensive RL training. In this paper, we propose OVExp, a training-efficient framework for open-vocabulary exploration. We make the first effort to demonstrate the generalization capabilities of semantic map-based goal prediction networks using Dense CLIP models. A major challenge is that preserving both precise point-wise object locations and generalizable visual representations in the semantic map leads to unaffordable training costs. To address this, we design a Cross-Modal Transfer on Semantic Mapping strategy which adapts an intriguing text-only training and transfer to multi-model semantic mapping and goals in test-time. Despite relying on text-based spatial layouts with limited objects, OVExp demonstrates robust generalization to unseentargets on established ObjectNav benchmarks.

📄 PDF Abstract BibTeX arXiv:2407.09016

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation

2024-09-22 · Naoki Yokoyama, Ram Ramrakhya, Abhishek Das, Dhruv Batra 외

We present the Habitat-Matterport 3D Open Vocabulary Object Goal Navigation dataset (HM3D-OVON), a large-scale benchmark that broadens the scope and semantic range of prior Object Goal Navigation (ObjectNav) benchmarks. …

NavigateVisual Navigation

GoalSwarm: Multi-UAV Semantic Coordination for Open-Vocabulary Object Navigation

2026-03-13 · MoniJesu Wonders James, Amir Atef Habel, Aleksey Fedoseev, Dzmitry Tsetserokou arxiv

Cooperative visual semantic navigation is a foundational capability for aerial robot teams operating in unknown environments. However, achieving robust open-vocabulary object-goal navigation remains challenging due to th…

OVAL: Open-Vocabulary Augmented Memory Model for Lifelong Object Goal Navigation

2026-04-14 · Jiahua Pei, Yi Liu, Guoping Pan, Yuanhao Jiang 외 arxiv

Object Goal Navigation (ObjectNav) refers to an agent navigating to an object in an unseen environment, which is an ability often required in the accomplishment of complex tasks. While existing methods demonstrate profic…

LagMemo: Language 3D Gaussian Splatting Memory for Multi-modal Open-vocabulary Multi-goal Visual Navigation

2025-10-28 · Haotian Zhou, Xiaole Wang, He Li, Zhuo Qi 외 arxiv

Navigating to a designated goal using visual information is a fundamental capability for intelligent robots. To address the practical demands of multi-modal, open-vocabulary goal queries and multi-goal visual navigation,…

Visual Navigation

GoalVLM: VLM-driven Object Goal Navigation for Multi-Agent System

2026-03-18 · MoniJesu James, Amir Atef Habel, Aleksey Fedoseev, Dzmitry Tsetserokou arxiv

Object-goal navigation has traditionally been limited to ground robots with closed-set object vocabularies. Existing multi-agent approaches depend on precomputed probabilistic graphs tied to fixed category sets, precludi…

Spatial Reasoning