paper-with-me

Papers

Towards Identity Preserving Normal to Dysarthric Voice Conversion

2021-10-15 · Wen-Chin Huang, Bence Mark Halpern, Lester Phillip Violeta, Odette Scharenborg, Tomoki Toda

We present a voice conversion framework that converts normal speech into dysarthric speech while preserving the speaker identity. Such a framework is essential for (1) clinical decision making processes and alleviation of patient stress, (2) data augmentation for dysarthric speech recognition. This is an especially challenging task since the converted samples should capture the severity of dysarthric speech while being highly natural and possessing the speaker identity of the normal speaker. To this end, we adopted a two-stage framework, which consists of a sequence-to-sequence model and a nonparallel frame-wise model. Objective and subjective evaluations were conducted on the UASpeech dataset, and results showed that the method was able to yield reasonable naturalness and capture severity aspects of the pathological speech. On the other hand, the similarity to the normal source speaker's voice was limited and requires further improvements.

📄 PDF Abstract BibTeX arXiv:2110.08213

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDecision Makingspeech-recognitionSpeech RecognitionVoice Conversion

Similar Papers 제목 키워드 기반

A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice Conversion

2021-06-02 · Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Ching-Feng Liu 외

We propose a new paradigm for maintaining speaker identity in dysarthric voice conversion (DVC). The poor quality of dysarthric speech can be greatly improved by statistical VC, but as the normal speech utterances of a d…

Voice Conversion

DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model

2023-06-18 · Helin Wang, Thomas Thebaud, Jesus Villalba, Myra Sydnor 외

We present a novel typical-to-atypical voice conversion approach (DuTa-VC), which (i) can be trained with nonparallel data (ii) first introduces diffusion probabilistic model (iii) preserves the target speaker identity (…

Data AugmentationDecoderspeech-recognitionSpeech Recognition+1

Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech

2025-06-02 · Karl El Hajal, Enno Hermann, Sevada Hovsepyan, Mathew Magimai. -Doss

Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved A…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Rhythmspeech-recognition+2

Learning Explicit Prosody Models and Deep Speaker Embeddings for Atypical Voice Conversion

2020-11-03 · Disong Wang, Songxiang Liu, Lifa Sun, Xixin Wu 외

Though significant progress has been made for the voice conversion (VC) of typical speech, VC for atypical speech, e.g., dysarthric and second-language (L2) speech, remains a challenge, since it involves correcting for a…

speech-recognitionSpeech RecognitionVoice Conversion

Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages

2025-05-20 · Chin-Jou Li, Eunjung Yeo, Kwanghee Choi, Paula Andrea Pérez-Toro 외

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1