paper-with-me

Papers

Unsupervised speech enhancement with diffusion-based generative models

2023-09-19 · Berné Nortier, Mostafa Sadeghi, Romain Serizel

Recently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when generalising to unseen conditions. To address this issue, we introduce an alternative approach that operates in an unsupervised manner, leveraging the generative power of diffusion models. Specifically, in a training phase, a clean speech prior distribution is learnt in the short-time Fourier transform (STFT) domain using score-based diffusion models, allowing it to unconditionally generate clean speech from Gaussian noise. Then, we develop a posterior sampling methodology for speech enhancement by combining the learnt clean speech prior with a noise model for speech signal inference. The noise parameters are simultaneously learnt along with clean speech estimation through an iterative expectationmaximisation (EM) approach. To the best of our knowledge, this is the first work exploring diffusion-based generative models for unsupervised speech enhancement, demonstrating promising results compared to a recent variational auto-encoder (VAE)-based unsupervised approach and a state-of-the-art diffusion-based supervised method. It thus opens a new direction for future research in unsupervised speech enhancement.

📄 PDF Abstract BibTeX arXiv:2309.10450

Code (1)

joanne-b-nortier/udiffse 공식 구현 pytorch

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement

2025-07-03 · Mostafa Sadeghi, Jean-Eudes Ayilo, Romain Serizel, Xavier Alameda-Pineda arxiv

We explore unsupervised speech enhancement using diffusion models as expressive generative priors for clean speech. Existing approaches guide the reverse diffusion process using noisy speech through an approximate, noise…

Speech Enhancement

Diffusion-based Unsupervised Audio-visual Speech Enhancement

2024-10-04 · Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel, Xavier Alameda-Pineda

This paper proposes a new unsupervised audio-visual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. Firs…

Speech Enhancement

Diffusion-based speech enhancement with a weighted generative-supervised learning loss

2023-09-19 · Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel

Diffusion-based generative models have recently gained attention in speech enhancement (SE), providing an alternative to conventional supervised methods. These models transform clean speech training samples into Gaussian…

Speech Enhancement

A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

2023-03-23 · Chenshuang Zhang, Chaoning Zhang, Sheng Zheng, Mengchun Zhang 외

Generative AI has demonstrated impressive performance in various fields, among which speech synthesis is an interesting direction. With the diffusion model as the most popular generative model, numerous works have attemp…

Speech EnhancementSpeech SynthesisSurveytext-to-speech+2

Diffusion Buffer for Online Generative Speech Enhancement

2025-10-21 · Bunlong Lay, Rostislav Makarov, Simon Welker, Maris Hillemann 외 arxiv

Online Speech Enhancement was mainly reserved for predictive models. A key advantage of these models is that for an incoming signal frame from a stream of data, the model is called only once for enhancement. In contrast,…

Speech Enhancement