paper-with-me

Papers

Learning Explicit Prosody Models and Deep Speaker Embeddings for Atypical Voice Conversion

2020-11-03 · Disong Wang, Songxiang Liu, Lifa Sun, Xixin Wu, Xunying Liu, Helen Meng

Though significant progress has been made for the voice conversion (VC) of typical speech, VC for atypical speech, e.g., dysarthric and second-language (L2) speech, remains a challenge, since it involves correcting for atypical prosody while maintaining speaker identity. To address this issue, we propose a VC system with explicit prosodic modelling and deep speaker embedding (DSE) learning. First, a speech-encoder strives to extract robust phoneme embeddings from atypical speech. Second, a prosody corrector takes in phoneme embeddings to infer typical phoneme duration and pitch values. Third, a conversion model takes phoneme embeddings and typical prosody features as inputs to generate the converted speech, conditioned on the target DSE that is learned via speaker encoder or speaker adaptation. Extensive experiments demonstrate that speaker adaptation can achieve higher speaker similarity, and the speaker encoder based conversion model can greatly reduce dysarthric and non-native pronunciation patterns with improved speech intelligibility. A comparison of speech recognition results between the original dysarthric speech and converted speech show that absolute reduction of 47.6% character error rate (CER) and 29.3% word error rate (WER) can be achieved.

📄 PDF Abstract BibTeX arXiv:2011.01678

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionVoice Conversion

Similar Papers 제목 키워드 기반

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

2026-04-29 · Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews, Philipp Koehn 외 arxiv

To preserve or not to preserve prosody is a central question in voice anonymization. Prosody conveys meaning and affect, yet is tightly coupled with speaker identity. Existing methods either discard prosody for privacy o…

Towards zero-shot Text-based voice editing using acoustic context conditioning, utterance embeddings, and reference encoders

2022-10-28 · Jason Fong, Yun Wang, Prabhav Agrawal, Vimal Manohar 외

Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edited speech that is similar to the origin…

Speaker Verificationtext-to-speechText to Speech

Exploring VQ-VAE with Prosody Parameters for Speaker Anonymization

2024-09-24 · Sotheara Leang, Anderson Augusma, Eric Castelli, Frédérique Letué 외

Human speech conveys prosody, linguistic content, and speaker identity. This article investigates a novel speaker anonymization approach using an end-to-end network based on a Vector-Quantized Variational Auto-Encoder (V…

DecoderSpeaker anonymizationSpeaker Identification

Combining Automatic Speaker Verification and Prosody Analysis for Synthetic Speech Detection

2022-10-31 · Luigi Attorresi, Davide Salvi, Clara Borrelli, Paolo Bestagini 외

The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatic…

Audio CompressionFace SwappingRhythmSpeaker Verification+4

Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect

2025-05-27 · Jaya Narain, Vasudha Kowtha, Colin Lea, Lauren Tooley 외

Perceptual voice quality dimensions describe key characteristics of atypical speech and other speech modulations. Here we develop and evaluate voice quality models for seven voice and speech dimensions (intelligibility, …