paper-with-me

Papers

Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion

2024-07-15 · Jian Ma, Wenguan Wang, Yi Yang, Feng Zheng

Visual acoustic matching (VAM) is pivotal for enhancing the immersive experience, and the task of dereverberation is effective in improving audio intelligibility. Existing methods treat each task independently, overlooking the inherent reciprocity between them. Moreover, these methods depend on paired training data, which is challenging to acquire, impeding the utilization of extensive unpaired data. In this paper, we introduce MVSD, a mutual learning framework based on diffusion models. MVSD considers the two tasks symmetrically, exploiting the reciprocal relationship to facilitate learning from inverse tasks and overcome data scarcity. Furthermore, we employ the diffusion model as foundational conditional converters to circumvent the training instability and over-smoothing drawbacks of conventional GAN architectures. Specifically, MVSD employs two converters: one for VAM called reverberator and one for dereverberation called dereverberator. The dereverberator judges whether the reverberation audio generated by reverberator sounds like being in the conditional visual scenario, and vice versa. By forming a closed loop, these two converters can generate informative feedback signals to optimize the inverse tasks, even with easily acquired one-way unpaired data. Extensive experiments on two standard benchmarks, i.e., SoundSpaces-Speech and Acoustic AVSpeech, exhibit that our framework can improve the performance of the reverberator and dereverberator and better match specified visual scenarios.

📄 PDF Abstract BibTeX arXiv:2407.10373

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Novel-View Acoustic Synthesis from 3D Reconstructed Rooms

2023-10-23 · Byeongjoo Ahn, Karren Yang, Brian Hamilton, Jonathan Sheaffer 외

We investigate the benefit of combining blind audio recordings with 3D scene information for novel-view acoustic synthesis. Given audio recordings from 2-4 microphones and the 3D geometry and material of a scene containi…

3D geometrySound Source Localization

A Hybrid Model for Weakly-Supervised Speech Dereverberation

2025-02-06 · Louis Bahrman, Mathieu Fontaine, Gael Richard

This paper introduces a new training strategy to improve speech dereverberation systems using minimal acoustic information and reverberant (wet) speech. Most existing algorithms rely on paired dry/wet data, which is diff…

modelSpeech Dereverberation

Learning Audio-Visual Dereverberation

2021-06-14 · Changan Chen, Wei Sun, David Harwath, Kristen Grauman

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker IdentificationSpeech Enhancement+2

U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model

2025-07-17 · Louis Bahrman, Marius Rodrigues, Mathieu Fontaine, Gaël Richard arxiv

This paper explores the outcome of training state-of-the-art dereverberation models with supervision settings ranging from weakly-supervised to virtually unsupervised, relying solely on reverberant signals and an acousti…

A Composite T60 Regression and Classification Approach for Speech Dereverberation

2023-02-09 · Yuying Li, Yuchen Liu, Donald S. Williamson

Dereverberation is often performed directly on the reverberant audio signal, without knowledge of the acoustic environment. Reverberation time, T60, however, is an essential acoustic factor that reflects how reverberatio…

regressionSpeech Dereverberation