paper-with-me

홈 › Papers

Towards a Single ASR Model That Generalizes to Disordered Speech

2024-12-26 · Jimmy Tobin, Katrin Tomanek, Subhashini Venugopalan

This study investigates the impact of integrating a dataset of disordered speech recordings ($\sim$1,000 hours) into the fine-tuning of a near state-of-the-art ASR baseline system. Contrary to what one might expect, despite the data being less than 1% of the training data of the ASR system, we find a considerable improvement in disordered speech recognition accuracy. Specifically, we observe a 33% improvement on prompted speech, and a 26% improvement on a newly gathered spontaneous, conversational dataset of disordered speech. Importantly, there is no significant performance decline on standard speech recognition benchmarks. Further, we observe that the proposed tuning strategy helps close the gap between the baseline system and personalized models by 64% highlighting the significant progress as well as the room for improvement. Given the substantial benefits of our findings, this experiment suggests that from a fairness perspective, incorporating a small fraction of high quality disordered speech data in a training recipe is an easy step that could be done to make speech technology more accessible for users with speech disabilities.

📄 PDF Abstract BibTeX arXiv:2412.19315

Code (0)

등록된 구현이 없습니다.

Tasks

Fairnessspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Comparing Supervised Models And Learned Speech Representations For Classifying Intelligibility Of Disordered Speech On Selected Phrases

2021-07-08 · Subhashini Venugopalan, Joel Shor, Manoj Plakal, Jimmy Tobin 외

Automatic classification of disordered speech can provide an objective tool for identifying the presence and severity of speech impairment. Classification approaches can also help identify hard-to-recognize speech sample…

Task 2

Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech

2024-09-13 · Pan-Pan Jiang, Jimmy Tobin, Katrin Tomanek, Robert L. MacDonald 외

Project Euphonia, a Google initiative, is dedicated to improving automatic speech recognition (ASR) of disordered speech. A central objective of the project is to create a large, high-quality, and diverse speech corpus. …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1

Adversarial Data Augmentation for Disordered Speech Recognition

2021-08-02 · Zengrui Jin, Mengzhe Geng, Xurong Xie, Jianwei Yu 외

Automatic recognition of disordered speech remains a highly challenging task to date. The underlying neuro-motor conditions, often compounded with co-occurring physical disabilities, lead to the difficulty in collecting …

Data Augmentationspeech-recognitionSpeech Recognition

Investigation of Data Augmentation Techniques for Disordered Speech Recognition

2022-01-14 · Mengzhe Geng, Xurong Xie, Shansong Liu, Jianwei Yu 외

Disordered speech recognition is a highly challenging task. The underlying neuro-motor conditions of people with speech disorders, often compounded with co-occurring physical disabilities, lead to the difficulty in colle…

Data Augmentationspeech-recognitionSpeech Recognition

Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning

2024-12-25 · Chirag Nagpal, Subhashini Venugopalan, Jimmy Tobin, Marilyn Ladewig 외

We introduce a large language model (LLM) capable of processing speech inputs and show that tuning it further with reinforcement learning on human preference (RLHF) enables it to adapt better to disordered speech than tr…

Language ModelingLanguage ModellingLarge Language Modelreinforcement-learning+3