paper-with-me

홈 › Papers

SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding

2025-04-08 · CVPR 2025 1 · Mingfei Chen, Israel D. Gebru, Ishwarya Ananthabhotla, Christian Richardt, Dejan Markovic, Jake Sandakly, Steven Krenn, Todd Keebler, Eli Shlizerman, Alexander Richard

We introduce SoundVista, a method to generate the ambient sound of an arbitrary scene at novel viewpoints. Given a pre-acquired recording of the scene from sparsely distributed microphones, SoundVista can synthesize the sound of that scene from an unseen target viewpoint. The method learns the underlying acoustic transfer function that relates the signals acquired at the distributed microphones to the signal at the target viewpoint, using a limited number of known recordings. Unlike existing works, our method does not require constraints or prior knowledge of sound source details. Moreover, our method efficiently adapts to diverse room layouts, reference microphone configurations and unseen environments. To enable this, we introduce a visual-acoustic binding module that learns visual embeddings linked with local acoustic properties from panoramic RGB and depth data. We first leverage these embeddings to optimize the placement of reference microphones in any given scene. During synthesis, we leverage multiple embeddings extracted from reference locations to get adaptive weights for their contribution, conditioned on target viewpoint. We benchmark the task on both publicly available data and real-world settings. We demonstrate significant improvements over existing methods.

📄 PDF Abstract BibTeX arXiv:2504.05576

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modeling Ambient Scene Dynamics for Free-view Synthesis

2024-06-13 · Meng-Li Shih, Jia-Bin Huang, Changil Kim, Rajvi Shah 외

We introduce a novel method for dynamic free-view synthesis of an ambient scenes from a monocular capture bringing a immersive quality to the viewing experience. Our method builds upon the recent advancements in 3D Gauss…

3DGSGPUNovel View Synthesis

Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos

2024-06-13 · Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue 외

Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the v…

Audio GenerationRetrieval-augmented Generation

Novel-View Acoustic Synthesis

2023-01-20 · CVPR 2023 1 · Changan Chen, Alexander Richard, Roman Shapovalov, Vamsi Krishna Ithapu 외

We introduce the novel-view acoustic synthesis (NVAS) task: given the sight and sound observed at a source viewpoint, can we synthesize the sound of that scene from an unseen target viewpoint? We propose a neural renderi…

Neural RenderingNovel View Synthesis

Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

2021-05-03 · Yan-Bo Lin, Yu-Chiang Frank Wang

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural au…

Audio GenerationSelf-Supervised Learning

A Proposal for Foley Sound Synthesis Challenge

2022-07-21 · Keunwoo Choi, Sangshin Oh, Minsung Kang, Brian McFee

"Foley" refers to sound effects that are added to multimedia during post-production to enhance its perceived acoustic properties, e.g., by simulating the sounds of footsteps, ambient environmental sounds, or visible obje…