paper-with-me

홈 › Papers

Visual Speech Enhancement Without A Real Visual Stream

2020-12-20 · Sindhu B Hegde, K R Prajwal, Rudrabha Mukhopadhyay, Vinay Namboodiri, C. V. Jawahar

In this work, we re-think the task of speech enhancement in unconstrained real-world environments. Current state-of-the-art methods use only the audio stream and are limited in their performance in a wide range of real-world noises. Recent works using lip movements as additional cues improve the quality of generated speech over "audio-only" methods. But, these methods cannot be used for several applications where the visual stream is unreliable or completely absent. We propose a new paradigm for speech enhancement by exploiting recent breakthroughs in speech-driven lip synthesis. Using one such model as a teacher network, we train a robust student network to produce accurate lip movements that mask away the noise, thus acting as a "visual noise filter". The intelligibility of the speech enhanced by our pseudo-lip approach is comparable (< 3% difference) to the case of using real lips. This implies that we can exploit the advantages of using lip movements even in the absence of a real video stream. We rigorously evaluate our model using quantitative metrics as well as human evaluations. Additional ablation studies and a demo video on our website containing qualitative comparisons and results clearly illustrate the effectiveness of our approach. We provide a demo video which clearly illustrates the effectiveness of our proposed approach on our website: \url{http://cvit.iiit.ac.in/research/projects/cvit-projects/visual-speech-enhancement-without-a-real-visual-stream}. The code and models are also released for future research: \url{https://github.com/Sindhu-Hegde/pseudo-visual-speech-denoising}.

📄 PDF Abstract BibTeX arXiv:2012.10852

Code (1)

Sindhu-Hegde/pseudo-visual-speech-denoising 공식 구현 pytorch

Tasks

DenoisingSpeech DenoisingSpeech Enhancement

Similar Papers 제목 키워드 기반

Audio-Visual Speech Codecs: Rethinking Audio-Visual Speech Enhancement by Re-Synthesis

2022-03-31 · CVPR 2022 1 · Karren Yang, Dejan Markovic, Steven Krenn, Vasu Agrawal 외

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet…

Speech Enhancement

Real-Time System for Audio-Visual Target Speech Enhancement

2025-09-25 · T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang arxiv

We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…

Audio-Visual Speech RecognitionSpeech Enhancement

RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement

2024-07-10 · Honglie Chen, Rodrigo Mira, Stavros Petridis, Maja Pantic

In this paper, we aim to generate clean speech frame by frame from a live video stream and a noisy audio stream without relying on future inputs. To this end, we propose RT-LA-VocE, which completely re-designs every comp…

Speech Enhancement

Contextual Audio-Visual Switching For Speech Enhancement in Real-World Environments

2018-08-28 · Ahsan Adeel, Mandar Gogate, Amir Hussain

Human speech processing is inherently multimodal, where visual cues (lip movements) help to better understand the speech in noise. Lip-reading driven speech enhancement significantly outperforms benchmark audio-only appr…

Lip ReadingSpeech Enhancement

On the Role of Visual Cues in Audiovisual Speech Enhancement

2020-04-25 · Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald, Erik Marchi 외

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target spee…

Self-Supervised LearningSpeech Enhancement