paper-with-me

Papers

AdVerb: Visually Guided Audio Dereverberation

2023-08-23 · ICCV 2023 1 · Sanjoy Chowdhury, Sreyan Ghosh, Subhrajyoti Dasgupta, Anton Ratnarajah, Utkarsh Tyagi, Dinesh Manocha

We present AdVerb, a novel audio-visual dereverberation framework that uses visual cues in addition to the reverberant sound to estimate clean audio. Although audio-only dereverberation is a well-studied problem, our approach incorporates the complementary visual modality to perform audio dereverberation. Given an image of the environment where the reverberated sound signal has been recorded, AdVerb employs a novel geometry-aware cross-modal transformer architecture that captures scene geometry and audio-visual cross-modal relationship to generate a complex ideal ratio mask, which, when applied to the reverberant audio predicts the clean sound. The effectiveness of our method is demonstrated through extensive quantitative and qualitative evaluations. Our approach significantly outperforms traditional audio-only and audio-visual baselines on three downstream tasks: speech enhancement, speech recognition, and speaker verification, with relative improvements in the range of 18% - 82% on the LibriSpeech test-clean set. We also achieve highly satisfactory RT60 error scores on the AVSpeech dataset.

📄 PDF Abstract BibTeX arXiv:2308.12370

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Learning Audio-Visual Dereverberation

2021-06-14 · Changan Chen, Wei Sun, David Harwath, Kristen Grauman

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker IdentificationSpeech Enhancement+2

Blind Speech Separation and Dereverberation using Neural Beamforming

2021-03-24 · Lukas Pfeifenberger, Franz Pernkopf

In this paper, we present the Blind Speech Separation and Dereverberation (BSSD) network, which performs simultaneous speaker separation, dereverberation and speaker identification in a single neural network. Speaker sep…

Speaker IdentificationSpeaker SeparationSpeech SeparationTriplet

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

2026-05-01 · Akira Takahashi, Ryosuke Sawata, Shusuke Takahashi, Yuki Mitsufuji arxiv

Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model room-acoustic effects such as reverberation or room impulse responses (R…

Audio-visual multi-channel speech separation, dereverberation and recognition

2022-04-05 · Guinan Li, Jianwei Yu, Jiajun Deng, Xunying Liu 외

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverbera…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+2

Real-time Single-channel Dereverberation and Separation with Time-domainAudio Separation Network

2018-09-02 · ISCA Interspeech 2018 9 · Yi Luo, Nima Mesgarani

We investigate the recently proposed Time-domain Audio Sep-aration Network (TasNet) in the task of real-time single-channel speech dereverberation. Unlike systems that take time-frequency representation of the au…

DenoisingSpeech DereverberationSpeech Separation