paper-with-me

Papers

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

2026-05-01 · Akira Takahashi, Ryosuke Sawata, Shusuke Takahashi, Yuki Mitsufuji arxiv

Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model room-acoustic effects such as reverberation or room impulse responses (RIRs), and thus offer limited controllability over these effects. However, we hypothesize that such V2A models implicitly have semantic knowledge of the relationship between spatial audio and the corresponding vision cues. In this paper, we revisit a V2A model for the sake of the above, and propose the way to utilize the pretrained model as prior for physically grounded room-acoustic processing. Based on one of the state-of-the-art V2A models, MMAudio, we propose MMAudioReverbs that is a unified framework dealing with i) dereverberation and ii) room impulse response (RIR) estimation without network architectural modification, and fine-tuned on a small dataset. Experimental results showed that audio and visual cues respectively have advantage depending on the type of physical room acoustics. It implies that foundation V2A models can be used for physically grounded room-acoustic analysis.

📄 PDF Abstract BibTeX arXiv:2605.00431

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model

2025-07-17 · Louis Bahrman, Marius Rodrigues, Mathieu Fontaine, Gaël Richard arxiv

This paper explores the outcome of training state-of-the-art dereverberation models with supervision settings ranging from weakly-supervised to virtually unsupervised, relying solely on reverberant signals and an acousti…

Dereverberation of Autoregressive Envelopes for Far-field Speech Recognition

2021-08-12 · Anurenjan Purushothaman, Anirudh Sreeram, Rohit Kumar, Sriram Ganapathy

The task of speech recognition in far-field environments is adversely affected by the reverberant artifacts that elicit as the temporal smearing of the sub-band envelopes. In this paper, we develop a neural model for spe…

Speech Dereverberationspeech-recognitionSpeech Recognition

A Composite T60 Regression and Classification Approach for Speech Dereverberation

2023-02-09 · Yuying Li, Yuchen Liu, Donald S. Williamson

Dereverberation is often performed directly on the reverberant audio signal, without knowledge of the acoustic environment. Reverberation time, T60, however, is an essential acoustic factor that reflects how reverberatio…

regressionSpeech Dereverberation

Unsupervised Blind Joint Dereverberation and Room Acoustics Estimation with Diffusion Models

2024-08-14 · Jean-Marie Lemercier, Eloi Moliner, Simon Welker, Vesa Välimäki 외

This paper presents an unsupervised method for single-channel blind dereverberation and room impulse response (RIR) estimation, called BUDDy. The algorithm is rooted in Bayesian posterior sampling: it combines a likeliho…

Room Impulse Response (RIR)Speech Dereverberation

A Hybrid Model for Weakly-Supervised Speech Dereverberation

2025-02-06 · Louis Bahrman, Mathieu Fontaine, Gael Richard

This paper introduces a new training strategy to improve speech dereverberation systems using minimal acoustic information and reverberant (wet) speech. Most existing algorithms rely on paired dry/wet data, which is diff…

modelSpeech Dereverberation