paper-with-me

Papers

Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations

2018-04-09 · Ju-chieh Chou, Cheng-chieh Yeh, Hung-Yi Lee, Lin-shan Lee

Recently, cycle-consistent adversarial network (Cycle-GAN) has been successfully applied to voice conversion to a different speaker without parallel data, although in those approaches an individual model is needed for each target speaker. In this paper, we propose an adversarial learning framework for voice conversion, with which a single model can be trained to convert the voice to many different speakers, all without parallel data, by separating the speaker characteristics from the linguistic content in speech signals. An autoencoder is first trained to extract speaker-independent latent representations and speaker embedding separately using another auxiliary speaker classifier to regularize the latent representation. The decoder then takes the speaker-independent latent representation and the target speaker embedding as the input to generate the voice of the target speaker with the linguistic content of the source utterance. The quality of decoder output is further improved by patching with the residual signal produced by another pair of generator and discriminator. A target speaker set size of 20 was tested in the preliminary experiments, and very good voice quality was obtained. Conventional voice conversion metrics are reported. We also show that the speaker information has been properly reduced from the latent representations.

📄 PDF Abstract BibTeX arXiv:1804.02812

Code (3)

jjery2243542/voice_conversion 공식 구현 pytorch
BogiHsu/Voice-Conversion-PyTorch pytorch
arshd91/multitarget-voice-conversion-vctk pytorch

Tasks

DecoderVoice Conversion

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

MelGAN-VC: Voice Conversion and Audio Style Transfer on arbitrarily long samples using Spectrograms

2019-10-08 · Marco Pasini

Traditional voice conversion methods rely on parallel recordings of multiple speakers pronouncing the same sentences. For real-world applications however, parallel data is rarely available. We propose MelGAN-VC, a voice …

Generative Adversarial NetworkMusic Style TransferStyle TransferTranslation+1

VAW-GAN for Singing Voice Conversion with Non-parallel Training Data

2020-08-10 · Junchen Lu, Kun Zhou, Berrak Sisman, Haizhou Li

Singing voice conversion aims to convert singer's voice from source to target without changing singing content. Parallel training data is typically required for the training of singing voice conversion system, that is ho…

DecoderGenerative Adversarial NetworkVoice Conversion

Singing voice conversion with non-parallel data

2019-03-11 · Xin Chen, Wei Chu, Jinxi Guo, Ning Xu

Singing voice conversion is a task to convert a song sang by a source singer to the voice of a target singer. In this paper, we propose using a parallel data free, many-to-one voice conversion technique on singing voices…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Voice Conversion with Conditional SampleRNN

2018-08-24 · Cong Zhou, Michael Horgan, Vivek Kumar, Cristina Vasco 외

Here we present a novel approach to conditioning the SampleRNN generative model for voice conversion (VC). Conventional methods for VC modify the perceived speaker identity by converting between source and target acousti…

Voice Conversion

Semi-supervised voice conversion with amortized variational inference

2019-09-30 · Cory Stephenson, Gokce Keskin, Anil Thomas, Oguz H. Elibol

In this work we introduce a semi-supervised approach to the voice conversion problem, in which speech from a source speaker is converted into speech of a target speaker. The proposed method makes use of both parallel and…

Variational InferenceVoice Conversion