paper-with-me

Papers

Transfer Learning from Whisper for Microscopic Intelligibility Prediction

2024-04-02 · Paul Best, Santiago Cuervo, Ricard Marxer

Macroscopic intelligibility models predict the expected human word-error-rate for a given speech-in-noise stimulus. In contrast, microscopic intelligibility models aim to make fine-grained predictions about listeners' perception, e.g. predicting phonetic or lexical responses. State-of-the-art macroscopic models use transfer learning from large scale deep learning models for speech processing, whereas such methods have rarely been used for microscopic modeling. In this paper, we study the use of transfer learning from Whisper, a state-of-the-art deep learning model for automatic speech recognition, for microscopic intelligibility prediction at the level of lexical responses. Our method outperforms the considered baselines, even in a zero-shot setup, and yields a relative improvement of up to 66\% when fine-tuned to predict listeners' responses. Our results showcase the promise of large scale deep learning based methods for microscopic intelligibility prediction.

📄 PDF Abstract BibTeX arXiv:2404.01737

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionDeep LearningPredictionspeech-recognitionSpeech RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata

2023-09-18 · Ryandhimas E. Zezario, Fei Chen, Chiou-Shann Fuh, Hsin-Min Wang 외

Automated speech intelligibility assessment is pivotal for hearing aid (HA) development. In this paper, we present three novel methods to improve intelligibility prediction accuracy and introduce MBI-Net+, an enhanced ve…

Multi-Task LearningPredictionSelf-Supervised Learning

A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models

2024-09-16 · Ryandhimas E. Zezario, Sabato M. Siniscalchi, Hsin-Min Wang, Yu Tsao

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which use…

Automatic Speech RecognitionPrompt Engineeringspeech-recognitionSpeech Recognition

A Study on Incorporating Whisper for Robust Speech Assessment

2023-09-22 · Ryandhimas E. Zezario, Yu-Wen Chen, Szu-Wei Fu, Yu Tsao 외

This research introduces an enhanced version of the multi-objective speech assessment model--MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the e…

Self-Supervised Learning

Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models

2024-01-24 · Rhiannon Mogridge, George Close, Robert Sutherland, Thomas Hain 외

Neural networks have been successfully used for non-intrusive speech intelligibility prediction. Recently, the use of feature representations sourced from intermediate layers of pre-trained self-supervised and weakly-sup…

Decoder

Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs

2026-03-14 · Zijian Ling, Pingyi Hu, Xiuyong Gao, Xiaojing Ma 외 arxiv

Speech-driven large language models (LLMs) are increasingly accessed through speech interfaces, introducing new security risks via open acoustic channels. We present Sirens' Whisper (SWhisper), the first practical framew…