paper-with-me

Papers

DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model

2023-06-18 · Helin Wang, Thomas Thebaud, Jesus Villalba, Myra Sydnor, Becky Lammers, Najim Dehak, Laureano Moro-Velazquez

We present a novel typical-to-atypical voice conversion approach (DuTa-VC), which (i) can be trained with nonparallel data (ii) first introduces diffusion probabilistic model (iii) preserves the target speaker identity (iv) is aware of the phoneme duration of the target speaker. DuTa-VC consists of three parts: an encoder transforms the source mel-spectrogram into a duration-modified speaker-independent mel-spectrogram, a decoder performs the reverse diffusion to generate the target mel-spectrogram, and a vocoder is applied to reconstruct the waveform. Objective evaluations conducted on the UASpeech show that DuTa-VC is able to capture severity characteristics of dysarthric speech, reserves speaker identity, and significantly improves dysarthric speech recognition as a data augmentation. Subjective evaluations by two expert speech pathologists validate that DuTa-VC can preserve the severity and type of dysarthria of the target speakers in the synthesized speech.

📄 PDF Abstract BibTeX arXiv:2306.10588

Code (1)

wanghelin1997/duta-vc 공식 구현 pytorch

Tasks

Data AugmentationDecoderspeech-recognitionSpeech RecognitionVoice Conversion

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Learning Explicit Prosody Models and Deep Speaker Embeddings for Atypical Voice Conversion

2020-11-03 · Disong Wang, Songxiang Liu, Lifa Sun, Xixin Wu 외

Though significant progress has been made for the voice conversion (VC) of typical speech, VC for atypical speech, e.g., dysarthric and second-language (L2) speech, remains a challenge, since it involves correcting for a…

speech-recognitionSpeech RecognitionVoice Conversion

Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect

2025-05-27 · Jaya Narain, Vasudha Kowtha, Colin Lea, Lauren Tooley 외

Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice and speech dimensions (intelligibility, …

Improving fairness for spoken language understanding in atypical speech with Text-to-Speech

2023-11-16 · Helin Wang, Venkatesh Ravichandran, Milind Rao, Becky Lammers 외

Spoken language understanding (SLU) systems often exhibit suboptimal performance in processing atypical speech, typically caused by neurological conditions and motor impairments. Recent advancements in Text-to-Speech (TT…

Data AugmentationFairnessSpoken Language Understandingtext-to-speech+2

Unheard in the Digital Age: Rethinking AI Bias and Speech Diversity

2026-01-26 · Onyedikachi Hope Amaechi-Okorie, Branislav Radeljic arxiv

Speech remains one of the most visible yet overlooked vectors of inclusion and exclusion in contemporary society. While fluency is often equated with credibility and competence, individuals with atypical speech patterns …

Speech Recognition

Benchmarking VLMs' Reasoning About Persuasive Atypical Images

2024-09-16 · Sina Malakouti, Aysan Aghazadeh, Ashmit Khandelwal, Adriana Kovashka

Vision language models (VLMs) have shown strong zero-shot generalization across various tasks, especially when integrated with large language models (LLMs). However, their ability to comprehend rhetorical and persuasive …

BenchmarkingObject RecognitionZero-shot Generalization