paper-with-me

Papers

A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice Conversion

2021-06-02 · Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Ching-Feng Liu, Yu Tsao, Hsin-Min Wang, Tomoki Toda

We propose a new paradigm for maintaining speaker identity in dysarthric voice conversion (DVC). The poor quality of dysarthric speech can be greatly improved by statistical VC, but as the normal speech utterances of a dysarthria patient are nearly impossible to collect, previous work failed to recover the individuality of the patient. In light of this, we suggest a novel, two-stage approach for DVC, which is highly flexible in that no normal speech of the patient is required. First, a powerful parallel sequence-to-sequence model converts the input dysarthric speech into a normal speech of a reference speaker as an intermediate product, and a nonparallel, frame-wise VC model realized with a variational autoencoder then converts the speaker identity of the reference speech back to that of the patient while assumed to be capable of preserving the enhanced quality. We investigate several design options. Experimental evaluation results demonstrate the potential of our approach to improving the quality of the dysarthric speech while maintaining the speaker identity.

📄 PDF Abstract BibTeX arXiv:2106.01415

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

GhostVec: A New Threat to Speaker Privacy of End-to-End Speech Recognition System

2023-11-17 · Xiaojiao Chen, Sheng Li, Jiyi Li, Hao Huang 외

Speaker adaptation systems face privacy concerns, for such systems are trained on private datasets and often overfitting. This paper demonstrates that an attacker can extract speaker information by querying speaker-adapt…

DecoderPrivacy PreservingSpeaker Verificationspeech-recognition+1

Mask2Flow-TSE: Two-Stage Target Speaker Extraction with Masking and Flow Matching

2026-03-13 · Junwon Moon, Seungbeom Kim, Hansol Park, Hyunjin Choi 외 arxiv

Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech given a reference utterance. Existing masking-based approaches are lightweight and effective but suffer from an inability to syn…

Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling

2025-06-05 · David Palzer, Matthew Maciejewski, Eric Fosler-Lussier

In recent years, end-to-end approaches have made notable progress in addressing the challenge of speaker diarization, which involves segmenting and identifying speakers in multi-talker recordings. One such approach, Enco…

AttributeDecoderspeaker-diarizationSpeaker Diarization

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling

2025-06-14 · Hui Wang, Yifan Yang, Shujie Liu, Jinyu Li 외

Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design.…

text-to-speechText to Speech

Continual Learning for Personalized Co-speech Gesture Generation

2023-01-01 · ICCV 2023 1 · Chaitanya Ahuja, Pratik Joshi, Ryo Ishii, Louis-Philippe Morency

Co-speech gestures are a key channel of human communication, making them important for personalized chat agents to generate. In the past, gesture generation models assumed that data for each speaker is available all …

Continual LearningGesture Generation