paper-with-me

Papers

Unsupervised Domain Adaptation for Dysarthric Speech Detection via Domain Adversarial Training and Mutual Information Minimization

2021-06-18 · Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen, Xunying Liu, Helen Meng

Dysarthric speech detection (DSD) systems aim to detect characteristics of the neuromotor disorder from speech. Such systems are particularly susceptible to domain mismatch where the training and testing data come from the source and target domains respectively, but the two domains may differ in terms of speech stimuli, disease etiology, etc. It is hard to acquire labelled data in the target domain, due to high costs of annotating sizeable datasets. This paper makes a first attempt to formulate cross-domain DSD as an unsupervised domain adaptation (UDA) problem. We use labelled source-domain data and unlabelled target-domain data, and propose a multi-task learning strategy, including dysarthria presence classification (DPC), domain adversarial training (DAT) and mutual information minimization (MIM), which aim to learn dysarthria-discriminative and domain-invariant biomarker embeddings. Specifically, DPC helps biomarker embeddings capture critical indicators of dysarthria; DAT forces biomarker embeddings to be indistinguishable in source and target domains; and MIM further reduces the correlation between biomarker embeddings and domain-related cues. By treating the UASPEECH and TORGO corpora respectively as the source and target domains, experiments show that the incorporation of UDA attains absolute increases of 22.2% and 20.0% respectively in utterance-level weighted average recall and speaker-level accuracy.

📄 PDF Abstract BibTeX arXiv:2106.10127

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationMulti-Task LearningUnsupervised Domain Adaptation

Similar Papers 제목 키워드 기반

Interpretable Dysarthric Speaker Adaptation based on Optimal-Transport

2022-03-14 · Rosanna Turrisi, Leonardo Badino

This work addresses the mismatch problem between the distribution of training data (source) and testing data (target), in the challenging context of dysarthric speech recognition. We focus on Speaker Adaptation (SA) in c…

Domain Adaptationspeech-recognitionSpeech Recognition

Optimal Transport-based Adaptation in Dysarthric Speech Tasks

2021-04-06 · Rosanna Turrisi, Leonardo Badino

In many real-world applications, the mismatch between distributions of training data (source) and test data (target) significantly degrades the performance of machine learning algorithms. In speech data, causes of this m…

speech-recognitionSpeech Recognition

Recent Progress in the CUHK Dysarthric Speech Recognition System

2022-01-15 · Shansong Liu, Mengzhe Geng, Shoukang Hu, Xurong Xie 외

Despite the rapid progress of automatic speech recognition (ASR) technologies in the past few decades, recognition of disordered speech remains a highly challenging task to date. Disordered speech presents a wide spectru…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentation+3

Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition

2024-12-25 · Shujie Hu, Xurong Xie, Mengzhe Geng, Jiajun Deng 외

Data-intensive fine-tuning of speech foundation models (SFMs) to scarce and diverse dysarthric and elderly speech leads to data bias and poor generalization to unseen speakers. This paper proposes novel structured speake…

Attributespeech-recognitionSpeech Recognition

Interpretable Deep Learning Model for the Detection and Reconstruction of Dysarthric Speech

2019-07-10 · Daniel Korzekwa, Roberto Barra-Chicote, Bozena Kostek, Thomas Drugman 외

This paper proposed a novel approach for the detection and reconstruction of dysarthric speech. The encoder-decoder model factorizes speech into a low-dimensional latent space and encoding of the input text. We showed th…

Decoder