paper-with-me

Papers

Adversarial Data Augmentation for Disordered Speech Recognition

2021-08-02 · Zengrui Jin, Mengzhe Geng, Xurong Xie, Jianwei Yu, Shansong Liu, Xunying Liu, Helen Meng

Automatic recognition of disordered speech remains a highly challenging task to date. The underlying neuro-motor conditions, often compounded with co-occurring physical disabilities, lead to the difficulty in collecting large quantities of impaired speech required for ASR system development. To this end, data augmentation techniques play a vital role in current disordered speech recognition systems. In contrast to existing data augmentation techniques only modifying the speaking rate or overall shape of spectral contour, fine-grained spectro-temporal differences between disordered and normal speech are modelled using deep convolutional generative adversarial networks (DCGAN) during data augmentation to modify normal speech spectra into those closer to disordered speech. Experiments conducted on the UASpeech corpus suggest the proposed adversarial data augmentation approach consistently outperformed the baseline augmentation methods using tempo or speed perturbation on a state-of-the-art hybrid DNN system. An overall word error rate (WER) reduction up to 3.05\% (9.7\% relative) was obtained over the baseline system using no data augmentation. The final learning hidden unit contribution (LHUC) speaker adapted system using the best adversarial augmentation approach gives an overall WER of 25.89% on the UASpeech test set of 16 dysarthric speakers.

📄 PDF Abstract BibTeX arXiv:2108.00899

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Adversarial Data Augmentation Using VAE-GAN for Disordered Speech Recognition

2022-11-03 · Zengrui Jin, Xurong Xie, Mengzhe Geng, Tianzi Wang 외

Automatic recognition of disordered speech remains a highly challenging task to date. The underlying neuro-motor conditions, often compounded with co-occurring physical disabilities, lead to the difficulty in collecting …

Data AugmentationGenerative Adversarial Networkspeech-recognitionSpeech Recognition

Investigation of Data Augmentation Techniques for Disordered Speech Recognition

2022-01-14 · Mengzhe Geng, Xurong Xie, Shansong Liu, Jianwei Yu 외

Disordered speech recognition is a highly challenging task. The underlying neuro-motor conditions of people with speech disorders, often compounded with co-occurring physical disabilities, lead to the difficulty in colle…

Data Augmentationspeech-recognitionSpeech Recognition

Towards Automatic Data Augmentation for Disordered Speech Recognition

2023-12-14 · Zengrui Jin, Xurong Xie, Tianzi Wang, Mengzhe Geng 외

Automatic recognition of disordered speech remains a highly challenging task to date due to data scarcity. This paper presents a reinforcement learning (RL) based on-the-fly data augmentation approach for training state-…

Data AugmentationReinforcement Learning (RL)speech-recognitionSpeech Recognition

Spectro-Temporal Deep Features for Disordered Speech Assessment and Recognition

2022-01-14 · Mengzhe Geng, Shansong Liu, Jianwei Yu, Xurong Xie 외

Automatic recognition of disordered speech remains a highly challenging task to date. Sources of variability commonly found in normal speech including accent, age or gender, when further compounded with the underlying ca…

Data Augmentationspeech-recognitionSpeech Recognition

Recent Progress in the CUHK Dysarthric Speech Recognition System

2022-01-15 · Shansong Liu, Mengzhe Geng, Shoukang Hu, Xurong Xie 외

Despite the rapid progress of automatic speech recognition (ASR) technologies in the past few decades, recognition of disordered speech remains a highly challenging task to date. Disordered speech presents a wide spectru…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentation+3