paper-with-me

Papers

A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models

2024-09-16 · Ryandhimas E. Zezario, Sabato M. Siniscalchi, Hsin-Min Wang, Yu Tsao

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an audio-to-text module and evaluates the naturalness of text via targeted prompt engineering. We evaluate the assessment metrics predicted by GPT-4o and GPT-Whisper, examining their correlation with human-based quality and intelligibility assessments and the character error rate (CER) of automatic speech recognition. Experimental results show that GPT-4o alone is less effective for audio analysis, while GPT-Whisper achieves higher prediction accuracy, has moderate correlation with speech quality and intelligibility, and has higher correlation with CER. Compared to SpeechLMScore and DNSMOS, GPT-Whisper excels in intelligibility metrics, but performs slightly worse than SpeechLMScore in quality estimation. Furthermore, GPT-Whisper outperforms supervised non-intrusive models MOS-SSL and MTI-Net in Spearman's rank correlation for CER of Whisper. These findings validate GPT-Whisper's potential for zero-shot speech assessment without requiring additional training data.

📄 PDF Abstract BibTeX arXiv:2409.09914

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionPrompt Engineeringspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

HASA-net: A non-intrusive hearing-aid speech assessment network

2021-11-10 · Hsin-Tien Chiang, Yi-Chiao Wu, Cheng Yu, Tomoki Toda 외

Without the need of a clean reference, non-intrusive speech assessment methods have caught great attention for objective evaluations. Recently, deep neural network (DNN) models have been applied to build non-intrusive sp…

Speech Enhancement with Zero-Shot Model Selection

2020-12-17 · Ryandhimas E. Zezario, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao

Recent research on speech enhancement (SE) has seen the emergence of deep-learning-based methods. It is still a challenging task to determine the effective ways to increase the generalizability of SE under diverse test c…

Ensemble LearningmodelModel SelectionSpeech Enhancement+1

Multi-objective Non-intrusive Hearing-aid Speech Assessment Model

2023-11-15 · Hsin-Tien Chiang, Szu-Wei Fu, Hsin-Min Wang, Yu Tsao 외

Without the need for a clean reference, non-intrusive speech assessment methods have caught great attention for objective evaluations. While deep learning models have been used to develop non-intrusive speech assessment …

MetricNet: Towards Improved Modeling For Non-Intrusive Speech Quality Assessment

2021-04-02 · Meng Yu, Chunlei Zhang, Yong Xu, ShiXiong Zhang 외

The objective speech quality assessment is usually conducted by comparing received speech signal with its clean reference, while human beings are capable of evaluating the speech quality without any reference, such as in…

Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM

2018-08-16 · Szu-Wei Fu, Yu Tsao, Hsin-Te Hwang, Hsin-Min Wang

Nowadays, most of the objective speech quality assessment tools (e.g., perceptual evaluation of speech quality (PESQ)) are based on the comparison of the degraded/processed speech with its clean counterpart. The need of …

Speech Enhancement