paper-with-me

홈 › Papers

Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models

2023-10-11 · Bangguo Yu, Qihao Yuan, Kailai Li, Hamidreza Kasaei, Ming Cao

Visual target navigation is a critical capability for autonomous robots operating in unknown environments, particularly in human-robot interaction scenarios. While classical and learning-based methods have shown promise, most existing approaches lack common-sense reasoning and are typically designed for single-robot settings, leading to reduced efficiency and robustness in complex environments. To address these limitations, we introduce Co-NavGPT, a novel framework that integrates a Vision Language Model (VLM) as a global planner to enable common-sense multi-robot visual target navigation. Co-NavGPT aggregates sub-maps from multiple robots with diverse viewpoints into a unified global map, encoding robot states and frontier regions. The VLM uses this information to assign frontiers across the robots, facilitating coordinated and efficient exploration. Experiments on the Habitat-Matterport 3D (HM3D) demonstrate that Co-NavGPT outperforms existing baselines in terms of success rate and navigation efficiency, without requiring task-specific training. Ablation studies further confirm the importance of semantic priors from the VLM. We also validate the framework in real-world scenarios using quadrupedal robots. Supplementary video and code are available at: https://sites.google.com/view/co-navgpt2.

📄 PDF Abstract BibTeX arXiv:2310.07937

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningEfficient Exploration

Similar Papers 제목 키워드 기반

NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

2023-05-26 · Gengze Zhou, Yicong Hong, Qi Wu

Trained with an unprecedented scale of data, large language models (LLMs) like ChatGPT and GPT-4 exhibit the emergence of significant reasoning abilities from model scaling. Such a trend underscored the potential of trai…

Instruction FollowingVision and Language NavigationVisual Navigation

NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models

2024-07-17 · Gengze Zhou, Yicong Hong, Zun Wang, Xin Eric Wang 외

Capitalizing on the remarkable advancements in Large Language Models (LLMs), there is a burgeoning initiative to harness LLMs for instruction following robotic navigation. Such a trend underscores the potential of LLMs t…

Instruction FollowingVision and Language Navigation

When Semantics Connect the Swarm: LLM-Driven Fuzzy Control for Cooperative Multi-Robot Underwater Coverage

2025-11-02 · Jingzehua Xu, Weihang Zhang, Yangyang Li, Hongmiaoyi Zhang 외 arxiv

Underwater multi-robot cooperative coverage remains challenging due to partial observability, limited communication, environmental uncertainty, and the lack of access to global localization. To address these issues, this…

Semantic Communication

Distributed Visual-Inertial Cooperative Localization

2021-03-23 · Pengxiang Zhu, Patrick Geneva, Wei Ren, Guoquan Huang

In this paper we present a consistent and distributed state estimator for multi-robot cooperative localization (CL) which efficiently fuses environmental features and loop-closure constraints across time and robots. In p…

Decentralised and Cooperative Control of Multi-Robot Systems through Distributed Optimisation

2023-02-03 · Yi Dong, Zhongguo Li, Xingyu Zhao, Zhengtao Ding 외

Multi-robot cooperative control has gained extensive research interest due to its wide applications in civil, security, and military domains. This paper proposes a cooperative control algorithm for multi-robot systems wi…