paper-with-me

Papers

Aphasic Speech Recognition using a Mixture of Speech Intelligibility Experts

2020-08-25

Robust speech recognition is a key prerequisite for semantic feature extraction in automatic aphasic speech analysis. However, standard one-size-fits-all automatic speech recognition models perform poorly when applied to aphasic speech. One reason for this is the wide range of speech intelligibility due to different levels of severity (i.e., higher severity lends itself to less intelligible speech). To address this, we propose a novel acoustic model based on a mixture of experts (MoE), which handles the varying intelligibility stages present in aphasic speech by explicitly defining severity-based experts. At test time, the contribution of each expert is decided by estimating speech intelligibility with a speech intelligibility detector (SID). We show that our proposed approach significantly reduces phone error rates across all severity stages in aphasic speech compared to a baseline approach that does not incorporate severity information into the modeling process.

📄 PDF Abstract BibTeX arXiv:2008.10788

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Mixture-of-ExpertsRobust Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

AS-ASR: A Lightweight Framework for Aphasia-Specific Automatic Speech Recognition

2025-06-06 · Chen Bao, Chuanbing Huo, Qinyu Chen, Chang Gao

This paper proposes AS-ASR, a lightweight aphasia-specific speech recognition framework based on Whisper-tiny, tailored for low-resource deployment on edge devices. Our approach introduces a hybrid training strategy that…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Speech Data Augmentation for Improving Phoneme Transcriptions of Aphasic Speech Using Wav2Vec 2.0 for the PSST Challenge

2022-06-01 · RaPID (LREC) 2022 6 · Birger Moell, Jim O’Regan, Shivam Mehta, Ambika Kirkland 외

As part of the PSST challenge, we explore how data augmentations, data sources, and model size affect phoneme transcription accuracy on speech produced by individuals with aphasia. We evaluate model performance in terms …

Automatic Phoneme RecognitionData AugmentationPhoneme RecognitionRoom Impulse Response (RIR)

SLM-SS: Speech Language Model for Generative Speech Separation

2026-01-27 · Tianhua Li, Chenda Li, Wei Wang, Xin Zhou 외 arxiv

Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to maintain speech intelligibility in the s…

Speech RecognitionSpeech Separation

Automatic recognition and detection of aphasic natural speech

2024-08-26 · Mara Barberis, Pieter De Clercq, Bastiaan Tamm, Hugo Van hamme 외

Aphasia is a language disorder affecting one third of stroke patients. Current aphasia assessment does not consider natural speech due to the time consuming nature of manual transcriptions and a lack of knowledge on how …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Unsupervised Uncertainty Measures of Automatic Speech Recognition for Non-intrusive Speech Intelligibility Prediction

2022-04-08 · Zehai Tu, Ning Ma, Jon Barker

Non-intrusive intelligibility prediction is important for its application in realistic scenarios, where a clean reference signal is difficult to access. The construction of many non-intrusive predictors require either gr…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition