paper-with-me

홈 › Papers

Coded Speech Quality Measurement by a Non-Intrusive PESQ-DNN

2023-04-18 · Ziyi Xu, Ziyue Zhao, Tim Fingscheidt

Wideband codecs such as AMR-WB or EVS are widely used in (mobile) speech communication. Evaluation of coded speech quality is often performed subjectively by an absolute category rating (ACR) listening test. However, the ACR test is impractical for online monitoring of speech communication networks. Perceptual evaluation of speech quality (PESQ) is one of the widely used metrics instrumentally predicting the results of an ACR test. However, the PESQ algorithm requires an original reference signal, which is usually unavailable in network monitoring, thus limiting its applicability. NISQA is a new non-intrusive neural-network-based speech quality measure, focusing on super-wideband speech signals. In this work, however, we aim at predicting the well-known PESQ metric using a non-intrusive PESQ-DNN model. We illustrate the potential of this model by predicting the PESQ scores of wideband-coded speech obtained from AMR-WB or EVS codecs operating at different bitrates in noisy, tandeming, and error-prone transmission conditions. We compare our methods with the state-of-the-art network topologies of QualityNet, WaweNet, and DNSMOS -- all applied to PESQ prediction -- by measuring the mean absolute error (MAE) and the linear correlation coefficient (LCC). The proposed PESQ-DNN offers the best total MAE and LCC of 0.11 and 0.92, respectively, in conditions without frame loss, and still is best when including frame loss. Note that our model could be similarly used to non-intrusively predict POLQA or other (intrusive) metrics. Upon article acceptance, code will be provided at GitHub.

📄 PDF Abstract BibTeX arXiv:2304.09226

Code (1)

ifnspaml/PESQDNN 공식 구현 tf

Methods 이 논문이 사용한 방법론

Test 설명 없음
MAE 설명 없음
LCC Please enter a description about the method here

Similar Papers 제목 키워드 기반

Does a PESQNet (Loss) Require a Clean Reference Input? The Original PESQ Does, But ACR Listening Tests Don't

2022-05-04 · Ziyi Xu, Maximilian Strake, Tim Fingscheidt

Perceptual evaluation of speech quality (PESQ) requires a clean speech reference as input, but predicts the results from (reference-free) absolute category rating (ACR) tests. In this work, we train a fully convolutional…

Deep Noise Suppression Maximizing Non-Differentiable PESQ Mediated by a Non-Intrusive PESQNet

2021-11-06 · Ziyi Xu, Maximilian Strake, Tim Fingscheidt

Speech enhancement employing deep neural networks (DNNs) for denoising are called deep noise suppression (DNS). During training, DNS methods are typically trained with mean squared error (MSE) type loss functions, which …

DenoisingSpeech Enhancement

Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM

2018-08-16 · Szu-Wei Fu, Yu Tsao, Hsin-Te Hwang, Hsin-Min Wang

Nowadays, most of the objective speech quality assessment tools (e.g., perceptual evaluation of speech quality (PESQ)) are based on the comparison of the degraded/processed speech with its clean counterpart. The need of …

Speech Enhancement

Deep Noise Suppression With Non-Intrusive PESQNet Supervision Enabling the Use of Real Training Data

2021-03-31 · Ziyi Xu, Maximilian Strake, Tim Fingscheidt

Data-driven speech enhancement employing deep neural networks (DNNs) can provide state-of-the-art performance even in the presence of non-stationary noise. During the training process, most of the speech enhancement neur…

DenoisingSpeech Enhancement

A Study on Speech Assessment with Visual Cues

2025-06-11 · Shafique Ahmed, Ryandhimas E. Zezario, Nasir Saleem, Amir Hussain 외

Non-intrusive assessment of speech quality and intelligibility is essential when clean reference signals are unavailable. In this work, we propose a multimodal framework that integrates audio features and visual cues to …

Multi-Task Learning