paper-with-me

홈 › Papers

Multi-task single channel speech enhancement using speech presence probability as a secondary task training target

2020-11-15 · L. Wang, J. Zhu, I. Kodrasi

To cope with reverberation and noise in single channel acoustic scenarios, typical supervised deep neural network~(DNN)-based techniques learn a mapping from reverberant and noisy input features to a user-defined target. Commonly used targets are the desired signal magnitude, a time-frequency mask such as the Wiener gain, or the interference power spectral density and signal-to-interference ratio that can be used to compute a time-frequency mask. In this paper, we propose to incorporate multi-task learning in such DNN-based enhancement techniques by using speech presence probability (SPP) estimation as a secondary task assisting the target estimation in the main task. The advantage of multi-task learning lies in sharing domain-specific information between the two tasks (i.e., target and SPP estimation) and learning more generalizable and robust representations. To simultaneously learn both tasks, we propose to use the adaptive weighting method of losses derived from the homoscedastic uncertainty of tasks. Simulation results show that the dereverberation and noise reduction performance of a single-task DNN trained to directly estimate the Wiener gain is higher than the performance of single-task DNNs trained to estimate the desired signal magnitude, the interference power spectral density, or the signal-to-interference ratio. Incorporating the proposed multi-task learning scheme to jointly estimate the Wiener gain and the SPP increases the dereverberation and noise reduction further.

📄 PDF Abstract BibTeX arXiv:2011.07547

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task LearningSpeech Enhancement

Similar Papers 제목 키워드 기반

Student-Teacher Learning for BLSTM Mask-based Speech Enhancement

2018-03-27

Spectral mask estimation using bidirectional long short-term memory (BLSTM) neural networks has been widely used in various speech enhancement applications, and it has achieved great success when it is applied to multich…

Speech Enhancementspeech-recognitionSpeech Recognition

SRIB-LEAP submission to Far-field Multi-Channel Speech Enhancement Challenge for Video Conferencing

2021-06-24 · R G Prithvi Raj, Rohit Kumar, M K Jayesh, Anurenjan Purushothaman 외

This paper presents the details of the SRIB-LEAP submission to the ConferencingSpeech challenge 2021. The challenge involved the task of multi-channel speech enhancement to improve the quality of far field speech from mi…

Speech Enhancement

ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding

2022-07-19 · Yen-Ju Lu, Xuankai Chang, Chenda Li, Wangyou Zhang 외

This paper presents recent progress on integrating speech separation and enhancement (SSE) into the ESPnet toolkit. Compared with the previous ESPnet-SE work, numerous features have been added, including recent state-of-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Robust Speech RecognitionSpeech Enhancement+4

Closing the Gap Between Time-Domain Multi-Channel Speech Enhancement on Real and Simulation Conditions

2021-10-27 · Wangyou Zhang, Jing Shi, Chenda Li, Shinji Watanabe 외

The deep learning based time-domain models, e.g. Conv-TasNet, have shown great potential in both single-channel and multi-channel speech enhancement. However, many experiments on the time-domain speech enhancement model …

Speech Enhancementspeech-recognitionSpeech Recognition

INTERSPEECH 2021 ConferencingSpeech Challenge: Towards Far-field Multi-Channel Speech Enhancement for Video Conferencing

2021-04-02 · Wei Rao, Yihui Fu, Yanxin Hu, Xin Xu 외

The ConferencingSpeech 2021 challenge is proposed to stimulate research on far-field multi-channel speech enhancement for video conferencing. The challenge consists of two separate tasks: 1) Task 1 is multi-channel speec…

Speech EnhancementTask 2