paper-with-me

Papers

Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation

2024-07-26 · Shiyao Wang, Shiwan Zhao, Jiaming Zhou, Aobo Kong, Yong Qin

Dysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to new dysarthric speakers. Traditional speaker adaptation methodologies typically involve fine-tuning models for each speaker, but this strategy is cost-prohibitive and inconvenient for disabled users, requiring substantial data collection. To address this issue, we introduce a prototype-based approach that markedly improves DSR performance for unseen dysarthric speakers without additional fine-tuning. Our method employs a feature extractor trained with HuBERT to produce per-word prototypes that encapsulate the characteristics of previously unseen speakers. These prototypes serve as the basis for classification. Additionally, we incorporate supervised contrastive learning to refine feature extraction. By enhancing representation quality, we further improve DSR performance, enabling effective personalized DSR. We release our code at https://github.com/NKU-HLT/PB-DSR.

📄 PDF Abstract BibTeX arXiv:2407.18461

Code (1)

nku-hlt/pb-dsr 공식 구현 pytorch

Tasks

Contrastive Learningspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

The Effectiveness of Time Stretching for Enhancing Dysarthric Speech for Improved Dysarthric Speech Recognition

2022-01-13 · Luke Prananta, Bence Mark Halpern, Siyuan Feng, Odette Scharenborg

In this paper, we investigate several existing and a new state-of-the-art generative adversarial network-based (GAN) voice conversion method for enhancing dysarthric speech for improved dysarthric speech recognition. We …

Generative Adversarial NetworkPhoneme Recognitionspeech-recognitionSpeech Recognition+1

Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition

2024-12-25 · Shujie Hu, Xurong Xie, Mengzhe Geng, Jiajun Deng 외

Data-intensive fine-tuning of speech foundation models (SFMs) to scarce and diverse dysarthric and elderly speech leads to data bias and poor generalization to unseen speakers. This paper proposes novel structured speake…

Attributespeech-recognitionSpeech Recognition

Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR

2025-01-17 · Karl El Hajal, Enno Hermann, Ajinkya Kulkarni, Mathew Magimai. -Doss

Automatic speech recognition (ASR) systems are well known to perform poorly on dysarthric speech. Previous works have addressed this by speaking rate modification to reduce the mismatch with typical speech. Unfortunately…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Rhythmspeech-recognition+2

Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition

2025-01-25 · Satwinder Singh, Qianli Wang, Zihan Zhong, Clarion Mendes 외

In this paper, we present a speaker-independent dysarthric speech recognition system, with a focus on evaluating the recently released Speech Accessibility Project (SAP-1005) dataset, which includes speech data from indi…

speech-recognitionSpeech Recognition

Recent Progress in the CUHK Dysarthric Speech Recognition System

2022-01-15 · Shansong Liu, Mengzhe Geng, Shoukang Hu, Xurong Xie 외

Despite the rapid progress of automatic speech recognition (ASR) technologies in the past few decades, recognition of disordered speech remains a highly challenging task to date. Disordered speech presents a wide spectru…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentation+3