paper-with-me

홈 › Papers

Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation

2023-02-22 · Yuchen Hu, Chen Chen, Heqing Zou, Xionghu Zhong, Eng Siong Chng

Recent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they would degrade significantly when put under realistic noisy conditions, as the background noise could be mistaken for speaker's speech and thus interfere with the separated sources. To alleviate this problem, we propose a novel network to unify speech enhancement and separation with gradient modulation to improve noise-robustness. Specifically, we first build a unified network by combining speech enhancement (SE) and separation modules, with multi-task learning for optimization, where SE is supervised by parallel clean mixture to reduce noise for downstream speech separation. Furthermore, in order to avoid suppressing valid speaker information when reducing noise, we propose a gradient modulation (GM) strategy to harmonize the SE and SS tasks from optimization view. Experimental results show that our approach achieves the state-of-the-art on large-scale Libri2Mix- and Libri3Mix-noisy datasets, with SI-SNRi results of 16.0 dB and 15.8 dB respectively. Our code is available at GitHub.

📄 PDF Abstract BibTeX arXiv:2302.11131

Code (1)

yuchen005/unified-enhance-separation 공식 구현 pytorch

Tasks

Multi-Task LearningSpeech EnhancementSpeech Separationvalid

Similar Papers 제목 키워드 기반

A Single Speech Enhancement Model Unifying Dereverberation, Denoising, Speaker Counting, Separation, and Extraction

2023-10-12 · Kohei Saijo, Wangyou Zhang, Zhong-Qiu Wang, Shinji Watanabe 외

We propose a multi-task universal speech enhancement (MUSE) model that can perform five speech enhancement (SE) tasks: dereverberation, denoising, speech separation (SS), target speaker extraction (TSE), and speaker coun…

DenoisingSpeech EnhancementSpeech SeparationTarget Speaker Extraction

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

2025-10-23 · Haoyin Yan, Chengwei Liu, Shaofei Xue, Xiaotao Liang 외 arxiv

Neural audio codecs have largely promoted the application of language models (LMs) for speech applications. However, the effectiveness of autoregressive LM-based models in unifying speech enhancement (SE) tasks remains u…

Reinforcement LearningSpeech EnhancementSpeech Separation

Investigating self-supervised learning for speech enhancement and separation

2022-03-15 · Zili Huang, Shinji Watanabe, Shu-wen Yang, Paola Garcia 외

Speech enhancement and separation are two fundamental tasks for robust speech processing. Speech enhancement suppresses background noise while speech separation extracts target speech from interfering speakers. Despite a…

Self-Supervised LearningSpeech EnhancementSpeech Separation

A Modulation-Domain Loss for Neural-Network-based Real-time Speech Enhancement

2021-02-15

We describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a speaker identification task. The learne…

Speaker IdentificationSpeech DenoisingSpeech Enhancement

Improved Speech Enhancement with the Wave-U-Net

2018-11-27 · Craig Macartney, Tillman Weyde

We study the use of the Wave-U-Net architecture for speech enhancement, a model introduced by Stoller et al for the separation of music vocals and accompaniment. This end-to-end learning method for audio source separatio…

Audio Source SeparationSpeech Enhancementspeech-recognitionSpeech Recognition