paper-with-me

Papers

A Study on Incorporating Whisper for Robust Speech Assessment

2023-09-22 · Ryandhimas E. Zezario, Yu-Wen Chen, Szu-Wei Fu, Yu Tsao, Hsin-Min Wang, Chiou-Shann Fuh

This research introduces an enhanced version of the multi-objective speech assessment model--MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of Whisper in deploying a more robust speech assessment model. After that, we explore combining representations from Whisper and SSL models. The experimental results reveal that Whisper's embedding features can contribute to more accurate prediction performance. Moreover, combining the embedding features from Whisper and SSL models only leads to marginal improvement. As compared to intrusive methods, MOSA-Net, and other SSL-based speech assessment models, MOSA-Net+ yields notable improvements in estimating subjective quality and intelligibility scores across all evaluation metrics in Taiwan Mandarin Hearing In Noise test - Quality & Intelligibility (TMHINT-QI) dataset. To further validate its robustness, MOSA-Net+ was tested in the noisy-and-enhanced track of the VoiceMOS Challenge 2023, where it obtained the top-ranked performance among nine systems.

📄 PDF Abstract BibTeX arXiv:2309.12766

Code (1)

dhimasryan/tmhint-qi_voicemos2023 공식 구현

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models

2024-09-16 · Ryandhimas E. Zezario, Sabato M. Siniscalchi, Hsin-Min Wang, Yu Tsao

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which use…

Automatic Speech RecognitionPrompt Engineeringspeech-recognitionSpeech Recognition

Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment

2025-10-18 · Fu-An Chao, Bi-Cheng Yan, Berlin Chen arxiv

In this paper, we explore the untapped potential of Whisper, a well-established automatic speech recognition (ASR) foundation model, in the context of L2 spoken language assessment (SLA). Unlike prior studies that extrin…

Spoken Language UnderstandingSpeech Recognition

Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata

2023-09-18 · Ryandhimas E. Zezario, Fei Chen, Chiou-Shann Fuh, Hsin-Min Wang 외

Automated speech intelligibility assessment is pivotal for hearing aid (HA) development. In this paper, we present three novel methods to improve intelligibility prediction accuracy and introduce MBI-Net+, an enhanced ve…

Multi-Task LearningPredictionSelf-Supervised Learning

Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset

2024-12-25 · Neil Shah, Shirish Karande, Vineet Gandhi

Current Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intelligibility and fails to generalize well acr…

text-to-speechText to SpeechVoice Cloning

Incorporating Talker Identity Aids With Improving Speech Recognition in Adversarial Environments

2024-10-07 · Sagarika Alavilli, Annesya Banerjee, Gasser Elbanna, Annika Magaro

Current state-of-the-art speech recognition models are trained to map acoustic signals into sub-lexical units. While these models demonstrate superior performance, they remain vulnerable to out-of-distribution conditions…

Speaker Identificationspeech-recognitionSpeech Recognition