paper-with-me

Papers

Diffusion-based speech enhancement with a weighted generative-supervised learning loss

2023-09-19 · Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel

Diffusion-based generative models have recently gained attention in speech enhancement (SE), providing an alternative to conventional supervised methods. These models transform clean speech training samples into Gaussian noise centered at noisy speech, and subsequently learn a parameterized model to reverse this process, conditionally on noisy speech. Unlike supervised methods, generative-based SE approaches usually rely solely on an unsupervised loss, which may result in less efficient incorporation of conditioned noisy speech. To address this issue, we propose augmenting the original diffusion training objective with a mean squared error (MSE) loss, measuring the discrepancy between estimated enhanced speech and ground-truth clean speech at each reverse process iteration. Experimental results demonstrate the effectiveness of our proposed methodology.

📄 PDF Abstract BibTeX arXiv:2309.10457

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Unsupervised speech enhancement with diffusion-based generative models

2023-09-19 · Berné Nortier, Mostafa Sadeghi, Romain Serizel

Recently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when g…

Speech Enhancement

Diffusion-based Unsupervised Audio-visual Speech Enhancement

2024-10-04 · Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel, Xavier Alameda-Pineda

This paper proposes a new unsupervised audio-visual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. Firs…

Speech Enhancement

Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement

2025-07-03 · Mostafa Sadeghi, Jean-Eudes Ayilo, Romain Serizel, Xavier Alameda-Pineda arxiv

We explore unsupervised speech enhancement using diffusion models as expressive generative priors for clean speech. Existing approaches guide the reverse diffusion process using noisy speech through an approximate, noise…

Speech Enhancement

uSee: Unified Speech Enhancement and Editing with Conditional Diffusion Models

2023-10-02 · Muqiao Yang, Chunlei Zhang, Yong Xu, Zhongweiyang Xu 외

Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we…

DenoisingSelf-Supervised LearningSpeech DenoisingSpeech Enhancement

A weighted-variance variational autoencoder model for speech enhancement

2022-11-02 · Ali Golmakani, Mostafa Sadeghi, Xavier Alameda-Pineda, Romain Serizel

We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed …

Speech Enhancement