paper-with-me

홈 › Papers

Multi-goal Audio-visual Navigation using Sound Direction Map

2023-08-01 · Haru Kondoh, Asako Kanezaki

Over the past few years, there has been a great deal of research on navigation tasks in indoor environments using deep reinforcement learning agents. Most of these tasks use only visual information in the form of first-person images to navigate to a single goal. More recently, tasks that simultaneously use visual and auditory information to navigate to the sound source and even navigation tasks with multiple goals instead of one have been proposed. However, there has been no proposal for a generalized navigation task combining these two types of tasks and using both visual and auditory information in a situation where multiple sound sources are goals. In this paper, we propose a new framework for this generalized task: multi-goal audio-visual navigation. We first define the task in detail, and then we investigate the difficulty of the multi-goal audio-visual navigation task relative to the current navigation tasks by conducting experiments in various situations. The research shows that multi-goal audio-visual navigation has the difficulty of the implicit need to separate the sources of sound. Next, to mitigate the difficulties in this new task, we propose a method named sound direction map (SDM), which dynamically localizes multiple sound sources in a learning-based manner while making use of past memories. Experimental results show that the use of SDM significantly improves the performance of multiple baseline methods, regardless of the number of goals.

📄 PDF Abstract BibTeX arXiv:2308.00219

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement LearningNavigateVisual Navigation

Similar Papers 제목 키워드 기반

Semantic Audio-Visual Navigation

2020-12-21 · CVPR 2021 1 · Changan Chen, Ziad Al-Halah, Kristen Grauman

Recent work on audio-visual navigation assumes a constantly-sounding target and restricts the role of audio to signaling the target's position. We introduce semantic audio-visual navigation, where objects in the environm…

PositionVisual Navigation

Semantic Audio-Visual Navigation in Continuous Environments

2026-03-20 · Yichen Zeng, Hebaixu Wang, Meng Liu, Yu Zhou 외 arxiv

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on precomputed room impulse responses (RIRs) f…

Visual Navigation

Look, Listen, and Act: Towards Audio-Visual Embodied Navigation

2019-12-25 · Chuang Gan, Yiwei Zhang, Jiajun Wu, Boqing Gong 외

A crucial ability of mobile intelligent agents is to integrate the evidence from multiple sensory inputs in an environment and to make a sequence of actions to reach their goals. In this paper, we attempt to approach the…

Navigate

LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation

2026-07-04 · Rufeng Chen, Yue Chang, Zili Shao, Zhaofan Zhang 외 arxiv

Embodied navigation is moving toward long-horizon missions, yet existing long-horizon benchmarks are largely acoustically silent, and audio-visual navigation tasks typically focus on a single goal. We introduce LH-AVLN, …

Visual Navigation

AVLEN: Audio-Visual-Language Embodied Navigation in 3D Environments

2022-10-14 · Sudipta Paul, Amit K. Roy-Chowdhury, Anoop Cherian

Recent years have seen embodied visual navigation advance in two distinct directions: (i) in equipping the AI agent to follow natural language instructions, and (ii) in making the navigable world multimodal, e.g., audio-…

AI AgentHierarchical Reinforcement LearningNavigateVisual Navigation