Investigating Training Objectives for Generative Speech Enhancement
Generative speech enhancement has recently shown promising advancements in improving speech quality in noisy environments. Multiple diffusion-based frameworks exist, each employing distinct training objectives and learning techniques. This paper aims to explain the differences between these frameworks by focusing our investigation on score-based generative models and the Schr\"odinger bridge. We conduct a series of comprehensive experiments to compare their performance and highlight differing training behaviors. Furthermore, we propose a novel perceptual loss function tailored for the Schr\"odinger bridge framework, demonstrating enhanced performance and improved perceptual quality of the enhanced speech signals. All experimental code and pre-trained models are publicly available to facilitate further research and development in this domain.
Code (1)
Tasks
Speech EnhancementSimilar Papers 제목 키워드 기반
Investigating the Effects of Diffusion-based Conditional Generative Speech Models Used for Speech Enhancement on Dysarthric Speech
In this study, we aim to explore the effect of pre-trained conditional generative speech models for the first time on dysarthric speech due to Parkinson's disease recorded in an ideal/non-noisy condition. Considering one…
Speech EnhancementScale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
Recent speech enhancement models have shown impressive performance gains by scaling up model complexity and training data. However, the impact of dataset variability (e.g. text, language, speaker, and noise) has been und…
AttributeSpeech Enhancementtext-to-speechText to SpeechInvestigating the Impact of Speech Enhancement on Audio Deepfake Detection in Noisy Environments
Logical Access (LA) attacks, also known as audio deepfake attacks, use Text-to-Speech (TTS) or Voice Conversion (VC) methods to generate spoofed speech data. This can represent a serious threat to Automatic Speaker Verif…
Audio Deepfake DetectionSpeaker VerificationSpeech EnhancementVoice ConversionInvestigating the Design Space of Diffusion Models for Speech Enhancement
Diffusion models are a new class of generative models that have shown outstanding performance in image generation literature. As a consequence, studies have attempted to apply diffusion models to other tasks, such as spe…
Image GenerationSpeech EnhancementInvestigating the effect of residual and highway connections in speech enhancement models
Residual and skip connections play an important role in many current generative models. Although their theoretical and numerical advantages are understood, their role in speech enhancement systems has not been in…
DenoisingSpeech DenoisingSpeech Enhancement