paper-with-me

Papers

Interpretable Deep Learning Model for the Detection and Reconstruction of Dysarthric Speech

2019-07-10 · Daniel Korzekwa, Roberto Barra-Chicote, Bozena Kostek, Thomas Drugman, Mateusz Lajszczak

This paper proposed a novel approach for the detection and reconstruction of dysarthric speech. The encoder-decoder model factorizes speech into a low-dimensional latent space and encoding of the input text. We showed that the latent space conveys interpretable characteristics of dysarthria, such as intelligibility and fluency of speech. MUSHRA perceptual test demonstrated that the adaptation of the latent space let the model generate speech of improved fluency. The multi-task supervised approach for predicting both the probability of dysarthric speech and the mel-spectrogram helps improve the detection of dysarthria with higher accuracy. This is thanks to a low-dimensional latent space of the auto-encoder as opposed to directly predicting dysarthria from a highly dimensional mel-spectrogram.

📄 PDF Abstract BibTeX arXiv:1907.04743

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

Prototype-Based Disentanglement for Controllable Dysarthric Speech Synthesis

2026-02-09 · Haoshen Wang, Xueli Zhong, Bingbing Lin, Jia Huang 외 arxiv

Dysarthric speech exhibits high variability and limited labeled data, posing major challenges for both automatic speech recognition (ASR) and assistive speech technologies. Existing approaches rely on synthetic data augm…

Speech RecognitionData AugmentationSpeech Synthesis

CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction

2024-06-12 · Xueyuan Chen, Dongchao Yang, Dingdong Wang, Xixin Wu 외

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech. It still suffers from low speaker similarity and poor prosody naturalness. In this paper, we propose a multi-modal DSR model …

DecoderLanguage ModelingLanguage Modelling

Enhancement of Dysarthric Speech Reconstruction by Contrastive Learning

2024-10-05 · Keshvari Fatemeh, Mahdian Toroghi Rahil, Zareian Hassan

Dysarthric speech reconstruction is challenging due to its pathological sound patterns. Preserving speaker identity, especially without access to normal speech, is a key challenge. Our proposed approach uses contrastive …

Contrastive Learningspeech-recognitionSpeech Recognition

Interpretable Dysarthric Speaker Adaptation based on Optimal-Transport

2022-03-14 · Rosanna Turrisi, Leonardo Badino

This work addresses the mismatch problem between the distribution of training data (source) and testing data (target), in the challenging context of dysarthric speech recognition. We focus on Speaker Adaptation (SA) in c…

Domain Adaptationspeech-recognitionSpeech Recognition

UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization

2024-01-26 · Yuejiao Wang, Xixin Wu, Disong Wang, Lingwei Meng 외

Dysarthric speech reconstruction (DSR) systems aim to automatically convert dysarthric speech into normal-sounding speech. The technology eases communication with speakers affected by the neuromotor disorder and enhances…

DecoderDomain AdaptationGenerative Adversarial NetworkRepresentation Learning+1