A Study on Incorporating Whisper for Robust Speech Assessment
This research introduces an enhanced version of the multi-objective speech assessment model--MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of Whisper in deploying a more robust speech assessment model. After that, we explore combining representations from Whisper and SSL models. The experimental results reveal that Whisper's embedding features can contribute to more accurate prediction performance. Moreover, combining the embedding features from Whisper and SSL models only leads to marginal improvement. As compared to intrusive methods, MOSA-Net, and other SSL-based speech assessment models, MOSA-Net+ yields notable improvements in estimating subjective quality and intelligibility scores across all evaluation metrics in Taiwan Mandarin Hearing In Noise test - Quality & Intelligibility (TMHINT-QI) dataset. To further validate its robustness, MOSA-Net+ was tested in the noisy-and-enhanced track of the VoiceMOS Challenge 2023, where it obtained the top-ranked performance among nine systems.
Code (1)
Tasks
Self-Supervised LearningSimilar Papers 제목 키워드 기반
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which use…
Automatic Speech RecognitionPrompt Engineeringspeech-recognitionSpeech RecognitionProbing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
In this paper, we explore the untapped potential of Whisper, a well-established automatic speech recognition (ASR) foundation model, in the context of L2 spoken language assessment (SLA). Unlike prior studies that extrin…
Spoken Language UnderstandingSpeech RecognitionNon-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
Automated speech intelligibility assessment is pivotal for hearing aid (HA) development. In this paper, we present three novel methods to improve intelligibility prediction accuracy and introduce MBI-Net+, an enhanced ve…
Multi-Task LearningPredictionSelf-Supervised LearningAdvancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset
Current Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intelligibility and fails to generalize well acr…
text-to-speechText to SpeechVoice CloningIncorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments
Current state-of-the-art speech recognition models are trained to map acoustic signals into sub-lexical units. While these models demonstrate superior performance, they remain vulnerable to out-of-distribution conditions…
Speaker Identificationspeech-recognitionSpeech Recognition