paper-with-me

홈 › Papers

Speech Enhancement with Score-Based Generative Models in the Complex STFT Domain

2022-03-31 · Simon Welker, Julius Richter, Timo Gerkmann

Score-based generative models (SGMs) have recently shown impressive results for difficult generative tasks such as the unconditional and conditional generation of natural images and audio signals. In this work, we extend these models to the complex short-time Fourier transform (STFT) domain, proposing a novel training task for speech enhancement using a complex-valued deep neural network. We derive this training task within the formalism of stochastic differential equations (SDEs), thereby enabling the use of predictor-corrector samplers. We provide alternative formulations inspired by previous publications on using generative diffusion models for speech enhancement, avoiding the need for any prior assumptions on the noise distribution and making the training task purely generative which, as we show, results in improved enhancement performance.

📄 PDF Abstract BibTeX arXiv:2203.17004

Code (2)

judiebig/DR-DiffuSE pytorch
sp-uhh/sgmse pytorch

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Efficient Transformer-based Speech Enhancement Using Long Frames and STFT Magnitudes

2022-06-23 · Danilo de Oliveira, Tal Peer, Timo Gerkmann

The SepFormer architecture shows very good results in speech separation. Like other learned-encoder models, it uses short frames, as they have been shown to obtain better performance in these cases. This results in a lar…

Speech EnhancementSpeech Separation

Unsupervised speech enhancement with diffusion-based generative models

2023-09-19 · Berné Nortier, Mostafa Sadeghi, Romain Serizel

Recently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when g…

Speech Enhancement

DiffPhase: Generative Diffusion-based STFT Phase Retrieval

2022-11-08 · Tal Peer, Simon Welker, Timo Gerkmann

Diffusion probabilistic models have been recently used in a variety of tasks, including speech enhancement and synthesis. As a generative approach, diffusion models have been shown to be especially suitable for imputatio…

ImputationRetrievalSpeech Enhancement

Single channel speech enhancement by colored spectrograms

2023-10-26 · Sania Gul, Muhammad Salman Khan, Muhammad Fazeel

Speech enhancement concerns the processes required to remove unwanted background sounds from the target speech to improve its quality and intelligibility. In this paper, a novel approach for single-channel speech enhance…

DenoisingGenerative Adversarial NetworkSpeech Enhancement

A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement

2024-01-19 · Yuewei Zhang, Huanbin Zou, Jie Zhu

Two-stage pipeline is popular in speech enhancement tasks due to its superiority over traditional single-stage methods. The current two-stage approaches usually enhance the magnitude spectrum in the first stage, and furt…

Speech Enhancement