paper-with-me

Papers

Deep Variational Generative Models for Audio-visual Speech Separation

2020-08-17 · Viet-Nhat Nguyen, Mostafa Sadeghi, Elisa Ricci, Xavier Alameda-Pineda

In this paper, we are interested in audio-visual speech separation given a single-channel audio recording as well as visual information (lips movements) associated with each speaker. We propose an unsupervised technique based on audio-visual generative modeling of clean speech. More specifically, during training, a latent variable generative model is learned from clean speech spectrograms using a variational auto-encoder (VAE). To better utilize the visual information, the posteriors of the latent variables are inferred from mixed speech (instead of clean speech) as well as the visual data. The visual modality also serves as a prior for latent variables, through a visual network. At test time, the learned generative model (both for speaker-independent and speaker-dependent scenarios) is combined with an unsupervised non-negative matrix factorization (NMF) variance model for background noise. All the latent variables and noise parameters are then estimated by a Monte Carlo expectation-maximization algorithm. Our experiments show that the proposed unsupervised VAE-based method yields better separation performance than NMF-based approaches as well as a supervised deep learning-based technique.

📄 PDF Abstract BibTeX arXiv:2008.07191

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Audio-visual Speech Enhancement Using Conditional Variational Auto-Encoders

2019-08-07 · Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin 외

Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech sign…

Speech Enhancement

Robust Unsupervised Audio-visual Speech Enhancement Using a Mixture of Variational Autoencoders

2019-11-10 · Mostafa Sadeghi, Xavier Alameda-Pineda

Recently, an audio-visual speech generative model based on variational autoencoder (VAE) has been proposed, which is combined with a nonnegative matrix factorization (NMF) model for noise variance to perform unsupervised…

Speech Enhancement

Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior

2025-09-17 · Yochai Yemini, Rami Ben-Ari, Sharon Gannot, Ethan Fetaya arxiv

In this paper, we address the problem of single-microphone speech separation in the presence of ambient noise. We propose a generative unsupervised technique that directly models both clean speech and structured noise co…

Speech Separation

Audio-visual speech enhancement with a deep Kalman filter generative model

2022-11-02 · Ali Golmakani, Mostafa Sadeghi, Romain Serizel

Deep latent variable generative models based on variational autoencoder (VAE) have shown promising performance for audiovisual speech enhancement (AVSE). The underlying idea is to learn a VAEbased audiovisual prior distr…

Speech Enhancement

Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss

2021-03-02 · Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka 외

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual spee…

Speech Separation