paper-with-me

Papers

Collaborative Visual Navigation

2021-07-02 · Haiyang Wang, Wenguan Wang, Xizhou Zhu, Jifeng Dai, LiWei Wang

As a fundamental problem for Artificial Intelligence, multi-agent system (MAS) is making rapid progress, mainly driven by multi-agent reinforcement learning (MARL) techniques. However, previous MARL methods largely focused on grid-world like or game environments; MAS in visually rich environments has remained less explored. To narrow this gap and emphasize the crucial role of perception in MAS, we propose a large-scale 3D dataset, CollaVN, for multi-agent visual navigation (MAVN). In CollaVN, multiple agents are entailed to cooperatively navigate across photo-realistic environments to reach target locations. Diverse MAVN variants are explored to make our problem more general. Moreover, a memory-augmented communication framework is proposed. Each agent is equipped with a private, external memory to persistently store communication information. This allows agents to make better use of their past communication information, enabling more efficient collaboration and robust long-term planning. In our experiments, several baselines and evaluation metrics are designed. We also empirically verify the efficacy of our proposed MARL approach across different MAVN task settings.

📄 PDF Abstract BibTeX arXiv:2107.01151

Code (1)

Haiyang-W/MAVN 공식 구현 pytorch

Tasks

Multi-agent Reinforcement LearningNavigateVisual Navigation

Methods 이 논문이 사용한 방법론

MAS This optimizer mix ADAM and SGD creating the MAS optimizer.

Similar Papers 제목 키워드 기반

Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation

2026-03-31 · Edoardo Zorzi, Francesco Taioli, Yiming Wang, Marco Cristani 외 arxiv

We propose Question-Asking Navigation (QAsk-Nav), the first reproducible benchmark for Collaborative Instance Object Navigation (CoIN) that enables an explicit, separate assessment of embodied navigation and collaborativ…

UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model

2025-11-24 · Changxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu 외 arxiv

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging. Recent research on enhancing language-gu…

Visual Reasoning

Advancing Audio-Visual Navigation Through Multi-Agent Collaboration in 3D Environments

2025-09-21 · Hailong Zhang, Yinfeng Yu, Liejun Wang, Fuchun Sun 외 arxiv

Intelligent agents often require collaborative strategies to achieve complex tasks beyond individual capabilities in real-world scenarios. While existing audio-visual navigation (AVN) research mainly focuses on single-ag…

Spatial ReasoningVisual Navigation

One-Shot Informed Robotic Visual Search in the Wild

2020-03-22 · Karim Koreitem, Florian Shkurti, Travis Manderson, Wei-Di Chang 외

We consider the task of underwater robot navigation for the purpose of collecting scientifically relevant video data for environmental monitoring. The majority of field robots that currently perform monitoring tasks in u…

NavigateRepresentation LearningRobot NavigationVisual Navigation

FollowMe: a Robust Person Following Framework Based on Re-Identification and Gestures

2023-11-21 · Federico Rollo, Andrea Zunino, Gennaro Raiola, Fabio Amadio 외

Human-robot interaction (HRI) has become a crucial enabler in houses and industries for facilitating operational flexibility. When it comes to mobile collaborative robots, this flexibility can be further increased due to…