paper-with-me

Papers

Voice Conversion for Stuttered Speech, Instruments, Unseen Languages and Textually Described Voices

2023-10-12 · Matthew Baas, Herman Kamper

Voice conversion aims to convert source speech into a target voice using recordings of the target speaker as a reference. Newer models are producing increasingly realistic output. But what happens when models are fed with non-standard data, such as speech from a user with a speech impairment? We investigate how a recent voice conversion model performs on non-standard downstream voice conversion tasks. We use a simple but robust approach called k-nearest neighbors voice conversion (kNN-VC). We look at four non-standard applications: stuttered voice conversion, cross-lingual voice conversion, musical instrument conversion, and text-to-voice conversion. The latter involves converting to a target voice specified through a text description, e.g. "a young man with a high-pitched voice". Compared to an established baseline, we find that kNN-VC retains high performance in stuttered and cross-lingual voice conversion. Results are more mixed for the musical instrument and text-to-voice conversion tasks. E.g., kNN-VC works well on some instruments like drums but not on others. Nevertheless, this shows that voice conversion models - and kNN-VC in particular - are increasingly applicable in a range of non-standard downstream tasks. But there are still limitations when samples are very far from the training distribution. Code, samples, trained models: https://rf5.github.io/sacair2023-knnvc-demo/.

📄 PDF Abstract BibTeX arXiv:2310.08104

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

Identifying Source Speakers for Voice Conversion based Spoofing Attacks on Speaker Verification Systems

2022-06-18 · Danwei Cai, Zexin Cai, Ming Li

An automatic speaker verification system aims to verify the speaker identity of a speech signal. However, a voice conversion system could manipulate a person's speech signal to make it sound like another speaker's voice …

Speaker IdentificationSpeaker VerificationVoice Conversion

StutterZero and StutterFormer: End-to-End Speech Conversion for Stuttering Transcription and Correction

2025-10-21 · Qianheng Xu arxiv

Over 70 million people worldwide experience stuttering, yet most automatic speech systems misinterpret disfluent utterances or fail to transcribe them accurately. Existing methods for stutter correction rely on handcraft…

Semantic SimilaritySpeech Recognition

Towards Robust Neural Vocoding for Speech Generation: A Survey

2019-12-05 · Po-chun Hsu, Chun-hsuan Wang, Andy T. Liu, Hung-Yi Lee

Recently, neural vocoders have been widely used in speech synthesis tasks, including text-to-speech and voice conversion. However, when encountering data distribution mismatch between training and inference, neural vocod…

Speech SynthesisSurveytext-to-speechText to Speech+1

HiFi-VC: High Quality ASR-Based Voice Conversion

2022-03-31 · A. Kashkin, I. Karpukhin, S. Shishkin

The goal of voice conversion (VC) is to convert input voice to match the target speaker's voice while keeping text and prosody intact. VC is usually used in entertainment and speaking-aid systems, as well as applied for …

speech-recognitionSpeech RecognitionVocal Bursts Intensity PredictionVoice Conversion

NVC-Net: End-to-End Adversarial Voice Conversion

2021-06-02 · Bac Nguyen, Fabien Cardinaux

Voice conversion has gained increasing popularity in many applications of speech synthesis. The idea is to change the voice identity from one speaker into another while keeping the linguistic content unchanged. Many voic…

GPUSpeech SynthesisVoice Conversion