paper-with-me

홈 › Papers

A Single Speech Enhancement Model Unifying Dereverberation, Denoising, Speaker Counting, Separation, and Extraction

2023-10-12 · Kohei Saijo, Wangyou Zhang, Zhong-Qiu Wang, Shinji Watanabe, Tetsunori Kobayashi, Tetsuji Ogawa

We propose a multi-task universal speech enhancement (MUSE) model that can perform five speech enhancement (SE) tasks: dereverberation, denoising, speech separation (SS), target speaker extraction (TSE), and speaker counting. This is achieved by integrating two modules into an SE model: 1) an internal separation module that does both speaker counting and separation; and 2) a TSE module that extracts the target speech from the internal separation outputs using target speaker cues. The model is trained to perform TSE if the target speaker cue is given and SS otherwise. By training the model to remove noise and reverberation, we allow the model to tackle the five tasks mentioned above with a single model, which has not been accomplished yet. Evaluation results demonstrate that the proposed MUSE model can successfully handle multiple tasks with a single model.

📄 PDF Abstract BibTeX arXiv:2310.08277

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeech EnhancementSpeech SeparationTarget Speaker Extraction

Similar Papers 제목 키워드 기반

To Dereverb Or Not to Dereverb? Perceptual Studies On Real-Time Dereverberation Targets

2022-06-16 · Jean-Marc Valin, Ritwik Giri, Shrikant Venkataramani, Umut Isik 외

In real life, room effect, also known as room reverberation, and the present background noise degrade the quality of speech. Recently, deep learning-based speech enhancement approaches have shown a lot of promise and sur…

DenoisingSpeech Enhancement

Phase-aware Single-stage Speech Denoising and Dereverberation with U-Net

2020-06-01 · Interspeech 2020 6

In this work, we tackle a denoising and dereverberation problem with a single-stage framework. Although denoising and dereverberation may be considered two separate challenging tasks, and thus, two modules are typically …

DenoisingSpeech DenoisingSpeech Enhancement

BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding

2024-08-13 · Alimjan Mattursun, Liejun Wang, Yinfeng Yu

Speech self-supervised learning (SSL) represents has achieved state-of-the-art (SOTA) performance in multiple downstream tasks. However, its application in speech enhancement (SE) tasks remains immature, offering opportu…

DenoisingSelf-Supervised LearningSpeech Enhancement

Schrödinger Bridge for Generative Speech Enhancement

2024-07-22 · Ante Jukić, Roman Korostik, Jagadeesh Balam, Boris Ginsburg

This paper proposes a generative speech enhancement model based on Schr\"odinger bridge (SB). The proposed model is employing a tractable SB to formulate a data-to-data process between the clean speech distribution and t…

DenoisingSpeech DenoisingSpeech DereverberationSpeech Enhancement

CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement

2022-09-22 · Sherif Abdulatif, Ruizhe Cao, Bin Yang

In this work, we further develop the conformer-based metric generative adversarial network (CMGAN) model for speech enhancement (SE) in the time-frequency (TF) domain. This paper builds on our previous work but takes a m…

Audio Super-ResolutionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoder+8