paper-with-me

Papers

Arabic Dysarthric Speech Recognition Using Adversarial and Signal-Based Augmentation

2023-06-07 · Massa Baali, Ibrahim Almakky, Shady Shehata, Fakhri Karray

Despite major advancements in Automatic Speech Recognition (ASR), the state-of-the-art ASR systems struggle to deal with impaired speech even with high-resource languages. In Arabic, this challenge gets amplified, with added complexities in collecting data from dysarthric speakers. In this paper, we aim to improve the performance of Arabic dysarthric automatic speech recognition through a multi-stage augmentation approach. To this effect, we first propose a signal-based approach to generate dysarthric Arabic speech from healthy Arabic speech by modifying its speed and tempo. We also propose a second stage Parallel Wave Generative (PWG) adversarial model that is trained on an English dysarthric dataset to capture language-independant dysarthric speech patterns and further augment the signal-adjusted speech samples. Furthermore, we propose a fine-tuning and text-correction strategies for Arabic Conformer at different dysarthric speech severity levels. Our fine-tuned Conformer achieved 18% Word Error Rate (WER) and 17.2% Character Error Rate (CER) on synthetically generated dysarthric speech from the Arabic commonvoice speech dataset. This shows significant WER improvement of 81.8% compared to the baseline model trained solely on healthy data. We perform further validation on real English dysarthric speech showing a WER improvement of 124% compared to the baseline trained only on healthy English LJSpeech dataset.

📄 PDF Abstract BibTeX arXiv:2306.04368

Code (1)

massabaali7/AR_Dysarthric 공식 구현

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

The Effectiveness of Time Stretching for Enhancing Dysarthric Speech for Improved Dysarthric Speech Recognition

2022-01-13 · Luke Prananta, Bence Mark Halpern, Siyuan Feng, Odette Scharenborg

In this paper, we investigate several existing and a new state-of-the-art generative adversarial network-based (GAN) voice conversion method for enhancing dysarthric speech for improved dysarthric speech recognition. We …

Generative Adversarial NetworkPhoneme Recognitionspeech-recognitionSpeech Recognition+1

Neural Model Reprogramming with Similarity Based Mapping for Low-Resource Spoken Command Recognition

2021-10-08 · Hao Yen, Pin-Jui Ku, Chao-Han Huck Yang, Hu Hu 외

In this study, we propose a novel adversarial reprogramming (AR) approach for low-resource spoken command recognition (SCR), and build an AR-SCR system. The AR procedure aims to modify the acoustic signals (from the targ…

Spoken Command RecognitionTransfer Learning

Improving Dysarthric Speech Intelligibility Using Cycle-consistent Adversarial Training

2020-01-10 · Seung Hee Yang, Minhwa Chung

Dysarthria is a motor speech impairment affecting millions of people. Dysarthric speech can be far less intelligible than those of non-dysarthric speakers, causing significant communication difficulties. The goal of our …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Domain Adversarial Neural Networks for Dysarthric Speech Recognition

2020-10-07 · Dominika Woszczyk, Stavros Petridis, David Millard

Speech recognition systems have improved dramatically over the last few years, however, their performance is significantly degraded for the cases of accented or impaired speech. This work explores domain adversarial neur…

Multi-Task Learningspeech-recognitionSpeech Recognition

Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition

2022-05-13 · Zengrui Jin, Mengzhe Geng, Jiajun Deng, Tianzi Wang 외

Despite the rapid progress of automatic speech recognition (ASR) technologies targeting normal speech, accurate recognition of dysarthric and elderly speech remains highly challenging tasks to date. It is difficult to co…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+1