paper-with-me

Papers

On using the UA-Speech and TORGO databases to validate automatic dysarthric speech classification approaches

2022-11-16 · Guilherme Schu, Parvaneh Janbakhshi, Ina Kodrasi

Although the UA-Speech and TORGO databases of control and dysarthric speech are invaluable resources made available to the research community with the objective of developing robust automatic speech recognition systems, they have also been used to validate a considerable number of automatic dysarthric speech classification approaches. Such approaches typically rely on the underlying assumption that recordings from control and dysarthric speakers are collected in the same noiseless environment using the same recording setup. In this paper, we show that this assumption is violated for the UA-Speech and TORGO databases. Using voice activity detection to extract speech and non-speech segments, we show that the majority of state-of-the-art dysarthria classification approaches achieve the same or a considerably better performance when using the non-speech segments of these databases than when using the speech segments. These results demonstrate that such approaches trained and validated on the UA-Speech and TORGO databases are potentially learning characteristics of the recording environment or setup rather than dysarthric speech characteristics. We hope that these results raise awareness in the research community about the importance of the quality of recordings when developing and evaluating automatic dysarthria classification approaches.

📄 PDF Abstract BibTeX arXiv:2211.08833

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech

2025-06-02 · Karl El Hajal, Enno Hermann, Sevada Hovsepyan, Mathew Magimai. -Doss

Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved A…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Rhythmspeech-recognition+2

Towards Personalized Federated Learning for Dysarthric Speech Recognition

2026-06-11 · Tao Zhong, Mengzhe Geng, Jiajun Deng, Shujie Hu 외 arxiv

Speech recognition is challenging for dysarthric speakers. While federated learning (FL)-based ASR can be an effective tool for protecting privacy, it suffers from heterogeneity issues caused by speaker variability. Forc…

Personalized Federated LearningSpeech Recognition

Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR

2025-01-17 · Karl El Hajal, Enno Hermann, Ajinkya Kulkarni, Mathew Magimai. -Doss

Automatic speech recognition (ASR) systems are well known to perform poorly on dysarthric speech. Previous works have addressed this by speaking rate modification to reduce the mismatch with typical speech. Unfortunately…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Rhythmspeech-recognition+2

Prototype-Based Disentanglement for Controllable Dysarthric Speech Synthesis

2026-02-09 · Haoshen Wang, Xueli Zhong, Bingbing Lin, Jia Huang 외 arxiv

Dysarthric speech exhibits high variability and limited labeled data, posing major challenges for both automatic speech recognition (ASR) and assistive speech technologies. Existing approaches rely on synthetic data augm…

Speech RecognitionData AugmentationSpeech Synthesis

DARS: Dysarthria-Aware Rhythm-Style Synthesis for ASR Enhancement

2026-03-02 · Minghui Wu, Xueling Liu, Jiahuan Fan, Haitao Tang 외 arxiv

Dysarthric speech exhibits abnormal prosody and significant speaker variability, presenting persistent challenges for automatic speech recognition (ASR). While text-to-speech (TTS)-based data augmentation has shown poten…

Speech RecognitionData Augmentation