paper-with-me

홈 › Papers

Does a PESQNet (Loss) Require a Clean Reference Input? The Original PESQ Does, But ACR Listening Tests Don't

2022-05-04 · Ziyi Xu, Maximilian Strake, Tim Fingscheidt

Perceptual evaluation of speech quality (PESQ) requires a clean speech reference as input, but predicts the results from (reference-free) absolute category rating (ACR) tests. In this work, we train a fully convolutional recurrent neural network (FCRN) as deep noise suppression (DNS) model, with either a non-intrusive or an intrusive PESQNet, where only the latter has access to a clean speech reference. The PESQNet is used as a mediator providing a perceptual loss during the DNS training to maximize the PESQ score of the enhanced speech signal. For the intrusive PESQNet, we investigate two topologies, called early-fusion (EF) and middle-fusion (MF) PESQNet, and compare to the non-intrusive PESQNet to evaluate and to quantify the benefits of employing a clean speech reference input during DNS training. Detailed analyses show that the DNS trained with the MF-intrusive PESQNet outperforms the Interspeech 2021 DNS Challenge baseline and the same DNS trained with an MSE loss by 0.23 and 0.12 PESQ points, respectively. Furthermore, we can show that only marginal benefits are obtained compared to the DNS trained with the non-intrusive PESQNet. Therefore, as ACR listening tests, the PESQNet does not necessarily require a clean speech reference input, opening the possibility of using real data for DNS training.

📄 PDF Abstract BibTeX arXiv:2205.02085

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deep Noise Suppression With Non-Intrusive PESQNet Supervision Enabling the Use of Real Training Data

2021-03-31 · Ziyi Xu, Maximilian Strake, Tim Fingscheidt

Data-driven speech enhancement employing deep neural networks (DNNs) can provide state-of-the-art performance even in the presence of non-stationary noise. During the training process, most of the speech enhancement neur…

DenoisingSpeech Enhancement

Deep Noise Suppression Maximizing Non-Differentiable PESQ Mediated by a Non-Intrusive PESQNet

2021-11-06 · Ziyi Xu, Maximilian Strake, Tim Fingscheidt

Speech enhancement employing deep neural networks (DNNs) for denoising are called deep noise suppression (DNS). During training, DNS methods are typically trained with mean squared error (MSE) type loss functions, which …

DenoisingSpeech Enhancement

Unsupervised Uncertainty Measures of Automatic Speech Recognition for Non-intrusive Speech Intelligibility Prediction

2022-04-08 · Zehai Tu, Ning Ma, Jon Barker

Non-intrusive intelligibility prediction is important for its application in realistic scenarios, where a clean reference signal is difficult to access. The construction of many non-intrusive predictors require either gr…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Noisy-target Training: A Training Strategy for DNN-based Speech Enhancement without Clean Speech

2021-01-21 · Takuya Fujimura, Yuma Koizumi, Kohei Yatabe, Ryoichi Miyazaki

Deep neural network (DNN)-based speech enhancement ordinarily requires clean speech signals as the training target. However, collecting clean signals is very costly because they must be recorded in a studio. This require…

Speech Enhancementspeech-recognitionSpeech Recognition

Employing Real Training Data for Deep Noise Suppression

2023-09-05 · Ziyi Xu, Marvin Sach, Jan Pirklbauer, Tim Fingscheidt

Most deep noise suppression (DNS) models are trained with reference-based losses requiring access to clean speech. However, sometimes an additive microphone model is insufficient for real-world applications. Accordingly,…