paper-with-me

Papers

GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning

2024-10-17 · Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel

Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial network (GAN) conditioned by the latent features of a discriminative model pre-trained for speech enhancement in low SNR scenarios. Our proposed method achieves superior performance compared to state-of-the-arts discriminative methods and also surpasses end-to-end (E2E) trained GAN models. We also investigate the impact of various configurations for conditioning the proposed GAN model with the discriminative model and assess their influence on enhancing speech quality

📄 PDF Abstract BibTeX arXiv:2410.13599

Code (0)

등록된 구현이 없습니다.

Tasks

Generative Adversarial NetworkSpeech Enhancement

Similar Papers 제목 키워드 기반

DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers

2025-04-13 · Heitor R. Guimarães, Jiaqi Su, Rithesh Kumar, Tiago H. Falk 외

Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity signals. While recent generative approa…

HallucinationSpeech Enhancement

A Systematic Comparison of Phonetic Aware Techniques for Speech Enhancement

2022-06-22 · Or Tal, Moshe Mandel, Felix Kreuk, Yossi Adi

Speech enhancement has seen great improvement in recent years using end-to-end neural networks. However, most models are agnostic to the spoken phonetic content. Recently, several studies suggested phonetic-aware speech …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Model OptimizationSelf-Supervised Learning+2

Disentanglement in a GAN for Unconditional Speech Synthesis

2023-07-04 · Matthew Baas, Herman Kamper

Can we develop a model that can synthesize realistic speech directly from a latent space, without explicit conditioning? Despite several efforts over the last decade, previous adversarial and diffusion-based approaches s…

DisentanglementGenerative Adversarial NetworkImage GenerationSpeaker Verification+3

Investigating the Design Space of Diffusion Models for Speech Enhancement

2023-12-07 · Philippe Gonzalez, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen 외

Diffusion models are a new class of generative models that have shown outstanding performance in image generation literature. As a consequence, studies have attempted to apply diffusion models to other tasks, such as spe…

Image GenerationSpeech Enhancement

Normalize Everything: A Preconditioned Magnitude-Preserving Architecture for Diffusion-Based Speech Enhancement

2025-05-08 · Julius Richter, Danilo de Oliveira, Timo Gerkmann

This paper presents a new framework for diffusion-based speech enhancement. Our method employs a Schroedinger bridge to transform the noisy speech distribution into the clean speech distribution. To stabilize and improve…

Image GenerationSpeech Enhancement