DDTSE: Discriminative Diffusion Model for Target Speech Extraction
Diffusion models have gained attention in speech enhancement tasks, providing an alternative to conventional discriminative methods. However, research on target speech extraction under multi-speaker noisy conditions remains relatively unexplored. Moreover, the superior quality of diffusion methods typically comes at the cost of slower inference speed. In this paper, we introduce the Discriminative Diffusion model for Target Speech Extraction (DDTSE). We apply the same forward process as diffusion models and utilize the reconstruction loss similar to discriminative methods. Furthermore, we devise a two-stage training strategy to emulate the inference process during model training. DDTSE not only works as a standalone system, but also can further improve the performance of discriminative models without additional retraining. Experimental results demonstrate that DDTSE not only achieves higher perceptual quality but also accelerates the inference process by 3 times compared to the conventional diffusion model.
Code (0)
등록된 구현이 없습니다.
Tasks
modelSpeech EnhancementSpeech ExtractionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Target Speech Extraction with Conditional Diffusion Model
Diffusion model-based speech enhancement has received increased attention since it can generate very natural enhanced signals and generalizes well to unseen conditions. Diffusion models have been explored for several sub…
DenoisingmodelSpeech DenoisingSpeech Enhancement+1MDDM: A Multi-view Discriminative Enhanced Diffusion-based Model for Speech Enhancement
With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which te…
Speech EnhancementEnhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
Target speech extraction (TSE) isolates the speech of a specific speaker from a multi-talker overlapped speech mixture. Most existing TSE models rely on discriminative methods, typically predicting a time-frequency spect…
Speech ExtractionInformed Source Extraction With Application to Acoustic Echo Reduction
Informed speaker extraction aims to extract a target speech signal from a mixture of sources given prior knowledge about the desired speaker. Recent deep learning-based methods leverage a speaker discriminative model tha…
Acoustic echo cancellationSpeaker SeparationSoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary audio (a.k.a. cue audio). Although recent a…
Speech ExtractionSpeech Separation