paper-with-me

Papers

Novel-View Acoustic Synthesis from 3D Reconstructed Rooms

2023-10-23 · Byeongjoo Ahn, Karren Yang, Brian Hamilton, Jonathan Sheaffer, Anurag Ranjan, Miguel Sarabia, Oncel Tuzel, Jen-Hao Rick Chang

We investigate the benefit of combining blind audio recordings with 3D scene information for novel-view acoustic synthesis. Given audio recordings from 2-4 microphones and the 3D geometry and material of a scene containing multiple unknown sound sources, we estimate the sound anywhere in the scene. We identify the main challenges of novel-view acoustic synthesis as sound source localization, separation, and dereverberation. While naively training an end-to-end network fails to produce high-quality results, we show that incorporating room impulse responses (RIRs) derived from 3D reconstructed rooms enables the same network to jointly tackle these tasks. Our method outperforms existing methods designed for the individual tasks, demonstrating its effectiveness at utilizing 3D visual information. In a simulated study on the Matterport3D-NVAS dataset, our model achieves near-perfect accuracy on source localization, a PSNR of 26.44dB and a SDR of 14.23dB for source separation and dereverberation, resulting in a PSNR of 25.55 dB and a SDR of 14.20 dB on novel-view acoustic synthesis. We release our code and model on our project website at https://github.com/apple/ml-nvas3d. Please wear headphones when listening to the results.

📄 PDF Abstract BibTeX arXiv:2310.15130

Code (1)

apple/ml-nvas3d 공식 구현 pytorch

Tasks

3D geometrySound Source Localization

Similar Papers 제목 키워드 기반

Few-shot Acoustic Synthesis with Multimodal Flow Matching

2026-03-19 · Amandine Brunetto arxiv

Generating audio that is acoustically consistent with a scene is essential for immersive virtual environments. Recent neural acoustic field methods enable spatially continuous sound rendering but remain scene-specific, r…

Generative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation

2025-01-22 · Jackie Lin, Georg Götz, Hermes Sampedro Llopis, Haukur Hafsteinsson 외

This paper describes the synthesis of the room acoustics challenge as a part of the generative data augmentation workshop at ICASSP 2025. The challenge defines a unique generative task that is designed to improve the qua…

Data AugmentationDiversity

Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark

2024-03-27 · CVPR 2024 1 · Ziyang Chen, Israel D. Gebru, Christian Richardt, Anurag Kumar 외

We present a new dataset called Real Acoustic Fields (RAF) that captures real acoustic room data from multiple modalities. The dataset includes high-quality and densely captured room impulse response data paired with mul…

Few-Shot LearningPose TrackingResponse Generation

A$^3$T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing

2022-03-18 · He Bai, Renjie Zheng, Junkun Chen, Xintong Li 외

Recently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation. However, all the above tasks are in the direction of spee…

Representation LearningSpeaker Verificationspeech-recognitionSpeech Recognition+5

Novel-View Acoustic Synthesis

2023-01-20 · CVPR 2023 1 · Changan Chen, Alexander Richard, Roman Shapovalov, Vamsi Krishna Ithapu 외

We introduce the novel-view acoustic synthesis (NVAS) task: given the sight and sound observed at a source viewpoint, can we synthesize the sound of that scene from an unseen target viewpoint? We propose a neural renderi…

Neural RenderingNovel View Synthesis