paper-with-me

Papers

A Composite Predictive-Generative Approach to Monaural Universal Speech Enhancement

2025-05-30 · Jie Zhang, Haoyin Yan, Xiaofei Li

It is promising to design a single model that can suppress various distortions and improve speech quality, i.e., universal speech enhancement (USE). Compared to supervised learning-based predictive methods, diffusion-based generative models have shown greater potential due to the generative capacities from degraded speech with severely damaged information. However, artifacts may be introduced in highly adverse conditions, and diffusion models often suffer from a heavy computational burden due to many steps for inference. In order to jointly leverage the superiority of prediction and generation and overcome the respective defects, in this work we propose a universal speech enhancement model called PGUSE by combining predictive and generative modeling. Our model consists of two branches: the predictive branch directly predicts clean samples from degraded signals, while the generative branch optimizes the denoising objective of diffusion models. We utilize the output fusion and truncated diffusion scheme to effectively integrate predictive and generative modeling, where the former directly combines results from both branches and the latter modifies the reverse diffusion process with initial estimates from the predictive branch. Extensive experiments on several datasets verify the superiority of the proposed model over state-of-the-art baselines, demonstrating the complementarity and benefits of combining predictive and generative modeling.

📄 PDF Abstract BibTeX arXiv:2505.24576

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

SE Territory: Monaural Speech Enhancement Meets the Fixed Virtual Perceptual Space Mapping

2023-11-03 · Xinmeng Xu, Yuhong Yang, Weiping tu

Monaural speech enhancement has achieved remarkable progress recently. However, its performance has been constrained by the limited spatial cues available at a single microphone. To overcome this limitation, we introduce…

Multi-Task LearningSpeech Enhancement

Injecting Spatial Information for Monaural Speech Enhancement via Knowledge Distillation

2022-12-02 · Xinmeng Xu, Weiping tu, Yuhong Yang

Monaural speech enhancement (SE) provides a versatile and cost-effective approach to SE tasks by utilizing recordings from a single microphone. However, the monaural SE lags performance behind multi-channel SE as the mon…

Knowledge DistillationSpeech Enhancement

Deep learning for monaural speech separation

2014-05-04 · ICASSP 2014 5 · Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, Paris Smaragdis

Monaural source separation is useful for many real-world applications though it is a challenging problem. In this paper, we study deep learning for monaural speech separation. We propose the joint optimization of the dee…

Deep LearningMulti-Speaker Source SeparationSpeech Separation

Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source Separation

2015-02-13 · Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, Paris Smaragdis

Monaural source separation is important for many real world applications. It is challenging because, with only a single channel of information available, without any constraints, an infinite number of solutions are possi…

DenoisingSpeech DenoisingSpeech Separation

Bridging the Gap Between Monaural Speech Enhancement and Recognition with Distortion-Independent Acoustic Modeling

2019-03-11 · Peidong Wang, Ke Tan, DeLiang Wang

Monaural speech enhancement has made dramatic advances since the introduction of deep learning a few years ago. Although enhanced speech has been demonstrated to have better intelligibility and quality for human listener…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1