paper-with-me

Papers

Diffusion-based Unsupervised Audio-visual Speech Enhancement

2024-10-04 · Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel, Xavier Alameda-Pineda

This paper proposes a new unsupervised audio-visual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion model is pre-trained on clean speech conditioned on corresponding video data to simulate the speech generative distribution. This pre-trained model is then paired with the NMF-based noise model to estimate clean speech iteratively. Specifically, a diffusion-based posterior sampling approach is implemented within the reverse diffusion process, where after each iteration, a speech estimate is obtained and used to update the noise parameters. Experimental results confirm that the proposed AVSE approach not only outperforms its audio-only counterpart but also generalizes better than a recent supervised-generative AVSE method. Additionally, the new inference algorithm offers a better balance between inference speed and performance compared to the previous diffusion-based method. Code and demo available at: https://jeaneudesayilo.github.io/fast_UdiffSE

📄 PDF Abstract BibTeX arXiv:2410.05301

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

2026-06-16 · Colombe Mboungou, Mostafa Sadeghi, Jean-Eudes Ayilo, Romain Serizel arxiv

Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments. Recent work introduced diffusion-based unsupervised AVSE, where a speech diffusion model conditio…

Speech Enhancement

Robust Unsupervised Audio-visual Speech Enhancement Using a Mixture of Variational Autoencoders

2019-11-10 · Mostafa Sadeghi, Xavier Alameda-Pineda

Recently, an audio-visual speech generative model based on variational autoencoder (VAE) has been proposed, which is combined with a nonnegative matrix factorization (NMF) model for noise variance to perform unsupervised…

Speech Enhancement

AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement

2023-09-14 · Ju-chieh Chou, Chung-Ming Chien, Karen Livescu

Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data available; most audio-visual datasets are co…

ResynthesisSpeech Enhancement

Audio-visual Speech Enhancement Using Conditional Variational Auto-Encoders

2019-08-07 · Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin 외

Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech sign…

Speech Enhancement

Audio-Visual Speech Enhancement with Score-Based Generative Models

2023-06-02 · Julius Richter, Simone Frintrop, Timo Gerkmann

This paper introduces an audio-visual speech enhancement system that leverages score-based generative models, also known as diffusion models, conditioned on visual information. In particular, we exploit audio-visual embe…

Automatic Speech RecognitionLipreadingSpeech Enhancementspeech-recognition+1