paper-with-me

Papers

Towards Improved Zero-shot Voice Conversion with Conditional DSVAE

2022-05-11 · Jiachen Lian, Chunlei Zhang, Gopala Krishna Anumanchipalli, Dong Yu

Disentangling content and speaking style information is essential for zero-shot non-parallel voice conversion (VC). Our previous study investigated a novel framework with disentangled sequential variational autoencoder (DSVAE) as the backbone for information decomposition. We have demonstrated that simultaneous disentangling content embedding and speaker embedding from one utterance is feasible for zero-shot VC. In this study, we continue the direction by raising one concern about the prior distribution of content branch in the DSVAE baseline. We find the random initialized prior distribution will force the content embedding to reduce the phonetic-structure information during the learning process, which is not a desired property. Here, we seek to achieve a better content embedding with more phonetic information preserved. We propose conditional DSVAE, a new model that enables content bias as a condition to the prior modeling and reshapes the content embedding sampled from the posterior distribution. In our experiment on the VCTK dataset, we demonstrate that content embeddings derived from the conditional DSVAE overcome the randomness and achieve a much better phoneme classification accuracy, a stabilized vocalization and a better zero-shot VC performance compared with the competitive DSVAE baseline.

📄 PDF Abstract BibTeX arXiv:2205.05227

Code (1)

jlian2/Improved-Voice-Conversion-with-Conditional-DSVAE

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching

2025-01-29 · Ha-Yeong Choi, JaeHan Park

Despite remarkable advancements in recent voice conversion (VC) systems, enhancing speaker similarity in zero-shot scenarios remains challenging. This challenge arises from the difficulty of generalizing and adapting spe…

DecoderIn-Context LearningVoice Conversion

AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss

2019-05-14 · Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang 외

Non-parallel many-to-many voice conversion, as well as zero-shot voice conversion, remain under-explored areas. Deep style transfer algorithms, such as generative adversarial networks (GAN) and conditional variational au…

Style TransferVoice Conversion

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

2025-05-22 · Advait Joglekar, Divyanshu Singh, Rooshil Rohit Bhatia, S. Umesh

Voice Conversion research in recent times has increasingly focused on improving the zero-shot capabilities of existing methods. Despite remarkable advancements, current architectures still tend to struggle in zero-shot c…

DecoderVoice Conversion

Zero-shot Voice Conversion via Self-supervised Prosody Representation Learning

2021-10-27 · Shijun Wang, Dimche Kostadinov, Damian Borth

Voice Conversion (VC) for unseen speakers, also known as zero-shot VC, is an attractive research topic as it enables a range of applications like voice customizing, animation production, and others. Recent work in this a…

DisentanglementRepresentation LearningVoice Conversion

F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder

2020-04-15 · Kaizhi Qian, Zeyu Jin, Mark Hasegawa-Johnson, Gautham J. Mysore

Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Many style-transfer-inspired methods such as generative adversarial networks (GANs) and variational autoencoders (…

Style TransferVoice Conversion