paper-with-me

Papers

Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

2023-05-18 · Hao Shi, Kazuki Shimada, Masato Hirano, Takashi Shibuya, Yuichiro Koyama, Zhi Zhong, Shusuke Takahashi, Tatsuya Kawahara, Yuki Mitsufuji

Diffusion-based generative speech enhancement (SE) has recently received attention, but reverse diffusion remains time-consuming. One solution is to initialize the reverse diffusion process with enhanced features estimated by a predictive SE system. However, the pipeline structure currently does not consider for a combined use of generative and predictive decoders. The predictive decoder allows us to use the further complementarity between predictive and diffusion-based generative SE. In this paper, we propose a unified system that use jointly generative and predictive decoders across two levels. The encoder encodes both generative and predictive information at the shared encoding level. At the decoded feature level, we fuse the two decoded features by generative and predictive decoders. Specifically, the two SE modules are fused in the initial and final diffusion steps: the initial fusion initializes the diffusion process with the predictive SE to improve convergence, and the final fusion combines the two complementary SE outputs to enhance SE performance. Experiments conducted on the Voice-Bank dataset demonstrate that incorporating predictive information leads to faster decoding and higher PESQ scores compared with other score-based diffusion SE (StoRM and SGMSE+).

📄 PDF Abstract BibTeX arXiv:2305.10734

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSpeech Enhancement

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

A Composite Predictive-Generative Approach to Monaural Universal Speech Enhancement

2025-05-30 · Jie Zhang, Haoyin Yan, Xiaofei Li

It is promising to design a single model that can suppress various distortions and improve speech quality, i.e., universal speech enhancement (USE). Compared to supervised learning-based predictive methods, diffusion-bas…

DenoisingSpeech Enhancement

Diffusion Buffer for Online Generative Speech Enhancement

2025-10-21 · Bunlong Lay, Rostislav Makarov, Simon Welker, Maris Hillemann 외 arxiv

Online Speech Enhancement was mainly reserved for predictive models. A key advantage of these models is that for an incoming signal frame from a stream of data, the model is called only once for enhancement. In contrast,…

Speech Enhancement

Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement

2025-05-07 · Rauf Nasretdinov, Roman Korostik, Ante Jukić

In this work, we investigate application of generative speech enhancement to improve the robustness of ASR models in noisy and reverberant conditions. We employ a recently-proposed speech enhancement model based on Schr\…

Robust Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation

2022-12-22 · Jean-Marie Lemercier, Julius Richter, Simon Welker, Timo Gerkmann

Diffusion models have shown a great ability at bridging the performance gap between predictive and generative approaches for speech enhancement. We have shown that they may even outperform their predictive counterparts f…

Speech DereverberationSpeech Enhancement

Single and Few-step Diffusion for Generative Speech Enhancement

2023-09-18 · Bunlong Lay, Jean-Marie Lemercier, Julius Richter, Timo Gerkmann

Diffusion models have shown promising results in speech enhancement, using a task-adapted diffusion process for the conditional generation of clean speech given a noisy mixture. However, at test time, the neural network …

DenoisingSpeech Enhancement