paper-with-me

홈 › Papers

Learning Audio-Visual Dereverberation

2021-06-14 · Changan Chen, Wei Sun, David Harwath, Kristen Grauman

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to learn to dereverberate speech from audio-visual observations. The visual environment surrounding a human speaker reveals important cues about the room geometry, materials, and speaker location, all of which influence the precise reverberation effects. We introduce Visually-Informed Dereverberation of Audio (VIDA), an end-to-end approach that learns to remove reverberation based on both the observed monaural sound and visual scene. In support of this new task, we develop a large-scale dataset SoundSpaces-Speech that uses realistic acoustic renderings of speech in real-world 3D scans of homes offering a variety of room acoustics. Demonstrating our approach on both simulated and real imagery for speech enhancement, speech recognition, and speaker identification, we show it achieves state-of-the-art performance and substantially improves over audio-only methods.

📄 PDF Abstract BibTeX arXiv:2106.07732

Code (1)

facebookresearch/learning-audio-visual-dereverberation 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker IdentificationSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

AdVerb: Visually Guided Audio Dereverberation

2023-08-23 · ICCV 2023 1 · Sanjoy Chowdhury, Sreyan Ghosh, Subhrajyoti Dasgupta, Anton Ratnarajah 외

We present AdVerb, a novel audio-visual dereverberation framework that uses visual cues in addition to the reverberant sound to estimate clean audio. Although audio-only dereverberation is a well-studied problem, our app…

Speaker VerificationSpeech Enhancementspeech-recognitionSpeech Recognition

Audio-visual multi-channel speech separation, dereverberation and recognition

2022-04-05 · Guinan Li, Jianwei Yu, Jiajun Deng, Xunying Liu 외

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverbera…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+2

Audio-visual End-to-end Multi-channel Speech Separation, Dereverberation and Recognition

2023-07-06 · Guinan Li, Jiajun Deng, Mengzhe Geng, Zengrui Jin 외

Accurate recognition of cocktail party speech containing overlapping speakers, noise and reverberation remains a highly challenging task to date. Motivated by the invariance of visual modality to acoustic signal corrupti…

Speech DereverberationSpeech EnhancementSpeech Separation

Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling

2025-10-29 · Jiarong Du, Zhan Jin, Peijun Yang, Juan Liu 외 arxiv

Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often exist complex acoustic environments, acc…

Speech Enhancement

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

2026-05-01 · Akira Takahashi, Ryosuke Sawata, Shusuke Takahashi, Yuki Mitsufuji arxiv

Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model room-acoustic effects such as reverberation or room impulse responses (R…