paper-with-me

Papers

Visual Acoustic Matching

2022-02-14 · CVPR 2022 1 · Changan Chen, Ruohan Gao, Paul Calamia, Kristen Grauman

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to re-synthesize the audio to match the target room acoustics as suggested by its visible geometry and materials. To address this novel task, we propose a cross-modal transformer model that uses audio-visual attention to inject visual properties into the audio and generate realistic audio output. In addition, we devise a self-supervised training objective that can learn acoustic matching from in-the-wild Web videos, despite their lack of acoustically mismatched audio. We demonstrate that our approach successfully translates human speech to a variety of real-world environments depicted in images, outperforming both traditional acoustic matching and more heavily supervised baselines.

📄 PDF Abstract BibTeX arXiv:2202.06875

Code (1)

see2sound/see2sound jax

Similar Papers 제목 키워드 기반

Self-Supervised Visual Acoustic Matching

2023-07-27 · NeurIPS 2023 11

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source a…

Diversity

A Self-Supervised Denoising Strategy for Underwater Acoustic Camera Imageries

2024-06-05 · Xiaoteng Zhou, Katsunori Mizuno, Yilong Zhang

In low-visibility marine environments characterized by turbidity and darkness, acoustic cameras serve as visual sensors capable of generating high-resolution 2D sonar images. However, acoustic camera images are interfere…

DenoisingImage Denoising

Conditional Flow Matching for Visually-Guided Acoustic Highlighting

2026-02-03 · Hugo Malard, Gael Le Lan, Daniel Wong, David Lou Alon 외 arxiv

Visually-guided acoustic highlighting seeks to rebalance audio in alignment with the accompanying video, creating a coherent audio-visual experience. While visual saliency and enhancement have been widely studied, acoust…

SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning

2022-06-16 · Changan Chen, Carl Schissler, Sanchit Garg, Philip Kobernik 외

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion

2024-07-15 · Jian Ma, Wenguan Wang, Yi Yang, Feng Zheng

Visual acoustic matching (VAM) is pivotal for enhancing the immersive experience, and the task of dereverberation is effective in improving audio intelligibility. Existing methods treat each task independently, overlooki…