paper-with-me

Papers

ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling

2024-04-24 · Arjun Somayazulu, Sagnik Majumder, Changan Chen, Kristen Grauman

An environment acoustic model represents how sound is transformed by the physical characteristics of an indoor environment, for any given source/receiver location. Traditional methods for constructing acoustic models involve expensive and time-consuming collection of large quantities of acoustic data at dense spatial locations in the space, or rely on privileged knowledge of scene geometry to intelligently select acoustic data sampling locations. We propose active acoustic sampling, a new task for efficiently building an environment acoustic model of an unmapped environment in which a mobile agent equipped with visual and acoustic sensors jointly constructs the environment acoustic model and the occupancy map on-the-fly. We introduce ActiveRIR, a reinforcement learning (RL) policy that leverages information from audio-visual sensor streams to guide agent navigation and determine optimal acoustic data sampling positions, yielding a high quality acoustic model of the environment from a minimal set of acoustic samples. We train our policy with a novel RL reward based on information gain in the environment acoustic model. Evaluating on diverse unseen indoor environments from a state-of-the-art acoustic simulation platform, ActiveRIR outperforms an array of methods--both traditional navigation agents based on spatial novelty and visual exploration as well as existing state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2404.16216

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

2026-07-28 · Kaneyoshi Hiratsuka, Benjamin Yen, Ryosuke Kojima arxiv

Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in…

Sound Source Localization

Deep Neural Object Analysis by Interactive Auditory Exploration with a Humanoid Robot

2018-07-03 · Manfred Eppe, Matthias Kerzel, Erik Strahl, Stefan Wermter

We present a novel approach for interactive auditory object analysis with a humanoid robot. The robot elicits sensory information by physically shaking visually indistinguishable plastic capsules. It gathers the resultin…

General ClassificationMaterial Classification

SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization

2026-01-19 · Naqcho Ali Mehdi, Mohammad Adeel, Aizaz Ali Larik arxiv

We present SoundPlot, an open-source framework for analyzing avian vocalizations through acoustic feature extraction, dimensionality reduction, and neural audio synthesis. The system transforms audio signals into a multi…

Dimensionality Reduction

Multi-Task Learning for Audio Visual Active Speaker Detection

2019-06-01 · The ActivityNet Large-Scale Activity Recognition Challenge Workshop, CVPR 2019 6 · Yuanhang Zhang, Jingyun Xiao, Shuang Yang, Shiguang Shan

This report describes the approach underlying our submission to the active speaker detection task (task B-2) of ActivityNet Challenge 2019. We introduce a new audio-visual model which builds upon a 3D-ResNet18 visual mod…

Active Speaker DetectionAudio-Visual Active Speaker DetectionLipreadingMulti-Task Learning+1

Online Audio-Visual Autoregressive Speaker Extraction

2025-06-02 · Zexu Pan, Wupeng Wang, Shengkui Zhao, Chong Zhang 외

This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less explored. We first propose a lightweight v…