paper-with-me

Papers

SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning

2022-06-16 · Changan Chen, Carl Schissler, Sanchit Garg, Philip Kobernik, Alexander Clegg, Paul Calamia, Dhruv Batra, Philip W Robinson, Kristen Grauman

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from arbitrary microphone locations. Together with existing 3D visual assets, it supports an array of audio-visual research tasks, such as audio-visual navigation, mapping, source localization and separation, and acoustic matching. Compared to existing resources, SoundSpaces 2.0 has the advantages of allowing continuous spatial sampling, generalization to novel environments, and configurable microphone and material properties. To our knowledge, this is the first geometry-based acoustic simulation that offers high fidelity and realism while also being fast enough to use for embodied learning. We showcase the simulator's properties and benchmark its performance against real-world audio measurements. In addition, we demonstrate two downstream tasks -- embodied navigation and far-field automatic speech recognition -- and highlight sim2real performance for the latter. SoundSpaces 2.0 is publicly available to facilitate wider research for perceptual systems that can both see and hear.

📄 PDF Abstract BibTeX arXiv:2206.08312

Code (2)

facebookresearch/rlr-audio-propagation 공식 구현
facebookresearch/sound-spaces 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionVisual Navigation

Similar Papers 제목 키워드 기반

SoundSpaces: Audio-Visual Navigation in 3D Environments

2019-12-24 · ECCV 2020 8 · Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari 외

Moving around in the world is naturally a multisensory experience, but today's embodied agents are deaf---restricted to solely their visual perception of the environment. We introduce audio-visual navigation for complex,…

Deep Reinforcement LearningNavigateReinforcement LearningVisual Navigation

Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction

2024-05-05 · Changan Chen, Jordi Ramos, Anshul Tomar, Kristen Grauman

Sim2real transfer has received increasing attention lately due to the success of learning robotic tasks in simulation end-to-end. While there has been a lot of progress in transferring vision-based navigation policies, t…

Data AugmentationNavigateVisual Navigation

Semantic Audio-Visual Navigation

2020-12-21 · CVPR 2021 1 · Changan Chen, Ziad Al-Halah, Kristen Grauman

Recent work on audio-visual navigation assumes a constantly-sounding target and restricts the role of audio to signaling the target's position. We introduce semantic audio-visual navigation, where objects in the environm…

PositionVisual Navigation

AudioWorldSim: Realistic Binaural Audio Datasets For World Models

2026-08-21 · Luis Vitor Zerkowski, Luiz Velho arxiv

This technical report presents AudioWorldSim, an open-source platform designed to generate realistic binaural audio datasets and advance research in audio-based machine learning, particularly world models. Built as a cus…

AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis

2023-02-04 · NeurIPS 2023 11

Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task -- real-world audio-visual scene synthes…

3D geometryAudio GenerationNeRF