paper-with-me

Papers

Prototype-Based Disentanglement for Controllable Dysarthric Speech Synthesis

2026-02-09 · Haoshen Wang, Xueli Zhong, Bingbing Lin, Jia Huang, Xingduo Pan, Shengxiang Liang, Nizhuan Wang, Wai Ting Siok arxiv

Dysarthric speech exhibits high variability and limited labeled data, posing major challenges for both automatic speech recognition (ASR) and assistive speech technologies. Existing approaches rely on synthetic data augmentation or speech reconstruction, yet often entangle speaker identity with pathological articulation, limiting controllability and robustness. In this paper, we propose ProtoDisent-TTS, a prototype-based disentanglement TTS framework built on a pre-trained text-to-speech backbone that factorizes speaker timbre and dysarthric articulation within a unified latent space. A pathology prototype codebook provides interpretable and controllable representations of healthy and dysarthric speech patterns, while a dual-classifier objective with a gradient reversal layer enforces invariance of speaker embeddings to pathological attributes. Experiments on the TORGO dataset demonstrate that this design enables bidirectional transformation between healthy and dysarthric speech, leading to consistent ASR performance gains and robust, speaker-aware speech reconstruction.

📄 PDF Abstract BibTeX arXiv:2602.08696

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionData AugmentationSpeech Synthesis

Similar Papers 제목 키워드 기반

Fairness in Dysarthric Speech Synthesis: Understanding Intrinsic Bias in Dysarthric Speech Cloning using F5-TTS

2025-08-07 · M Anuprabha, Krishna Gurugubelli, Anil Kumar Vuppala arxiv

Dysarthric speech poses significant challenges in developing assistive technologies, primarily due to the limited availability of data. Recent advances in neural speech synthesis, especially zero-shot voice cloning, faci…

Data AugmentationSpeech Synthesis

Accurate synthesis of Dysarthric Speech for ASR data augmentation

2023-08-16 · Mohammad Soleymanpour, Michael T. Johnson, Rahim Soleymanpour, Jeffrey Berry

Dysarthria is a motor speech disorder often characterized by reduced speech intelligibility through slow, uncoordinated control of speech production muscles. Automatic Speech recognition (ASR) systems can help dysarthric…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+2

Augmenting Dysarthric Speech Severity Assessment with MOS Supervision

2026-06-17 · Kaimeng Jia, Minzhu Tu, Zengrui Jin, Siyin Wang 외 arxiv

Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness. Automatic utterance-level assessment of dysarthric speech can support scalable speech monitoring and therapy-related anal…

Speech Synthesis

Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation

2024-07-26 · Shiyao Wang, Shiwan Zhao, Jiaming Zhou, Aobo Kong 외

Dysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to new dysarthric speakers. Traditional sp…

Contrastive Learningspeech-recognitionSpeech Recognition

DARS: Dysarthria-Aware Rhythm-Style Synthesis for ASR Enhancement

2026-03-02 · Minghui Wu, Xueling Liu, Jiahuan Fan, Haitao Tang 외 arxiv

Dysarthric speech exhibits abnormal prosody and significant speaker variability, presenting persistent challenges for automatic speech recognition (ASR). While text-to-speech (TTS)-based data augmentation has shown poten…

Speech RecognitionData Augmentation