paper-with-me

홈 › Papers

NoiseVC: Towards High Quality Zero-Shot Voice Conversion

2021-04-13 · Shijun Wang, Damian Borth

Voice conversion (VC) is a task that transforms voice from target audio to source without losing linguistic contents, it is challenging especially when source and target speakers are unseen during training (zero-shot VC). Previous approaches require a pre-trained model or linguistic data to do the zero-shot conversion. Meanwhile, VC models with Vector Quantization (VQ) or Instance Normalization (IN) are able to disentangle contents from audios and achieve successful conversions. However, disentanglement in these models highly relies on heavily constrained bottleneck layers, thus, the sound quality is drastically sacrificed. In this paper, we propose NoiseVC, an approach that can disentangle contents based on VQ and Contrastive Predictive Coding (CPC). Additionally, Noise Augmentation is performed to further enhance disentanglement capability. We conduct several experiments and demonstrate that NoiseVC has a strong disentanglement ability with a small sacrifice of quality.

📄 PDF Abstract BibTeX arXiv:2104.06074

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementQuantizationVocal Bursts Intensity PredictionVoice Conversion

Methods 이 논문이 사용한 방법론

InfoNCE 설명 없음
Instance Normalization Instance Normalization (also known as contrast normalization) is a normalization layer where: $$ y_{tijk} = \frac{x_{tijk} - \mu_{ti}}{\sqrt{\sigma_{ti}^2 +…
Contrastive Predictive Coding Contrastive Predictive Coding (CPC) learns self-supervised representations by predicting the future in latent space by using powerful autoregressive models. The model uses a…

Similar Papers 제목 키워드 기반

HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios

2025-11-11 · Bingsong Bai, Yizhong Geng, Fengping Wang, Cong Wang 외 arxiv

Zero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing methods model speaker timbre and vocal conte…

Audio Super-ResolutionVoice Conversion

StarGAN-ZSVC: Towards Zero-Shot Voice Conversion in Low-Resource Contexts

2021-05-31 · Matthew Baas, Herman Kamper

Voice conversion is the task of converting a spoken utterance from a source speaker so that it appears to be said by a different target speaker while retaining the linguistic content of the utterance. Recent advances hav…

Voice Conversion

YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone

2021-12-04 · Edresson Casanova, Julian Weber, Christopher Shulby, Arnaldo Candido Junior 외

YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual trai…

Speech SynthesisText-To-Speech SynthesisVoice ConversionVoice Similarity+2

VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching

2025-01-29 · Ha-Yeong Choi, JaeHan Park

Despite remarkable advancements in recent voice conversion (VC) systems, enhancing speaker similarity in zero-shot scenarios remains challenging. This challenge arises from the difficulty of generalizing and adapting spe…

DecoderIn-Context LearningVoice Conversion

Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model

2024-07-24 · Jan Lehečka, Zdeněk Hanzlíček, Jindřich Matoušek, Daniel Tihelka

In this paper, we experimented with the SpeechT5 model pre-trained on large-scale datasets. We pre-trained the foundation model from scratch and fine-tuned it on a large-scale robust multi-speaker text-to-speech (TTS) ta…

text-to-speechText to Speech