paper-with-me

Papers

Voice Conversion from Unaligned Corpora using Variational Autoencoding Wasserstein Generative Adversarial Networks

2017-04-04 · Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao, Hsin-Min Wang

Building a voice conversion (VC) system from non-parallel speech corpora is challenging but highly valuable in real application scenarios. In most situations, the source and the target speakers do not repeat the same texts or they may even speak different languages. In this case, one possible, although indirect, solution is to build a generative model for speech. Generative models focus on explaining the observations with latent variables instead of learning a pairwise transformation function, thereby bypassing the requirement of speech frame alignment. In this paper, we propose a non-parallel VC framework with a variational autoencoding Wasserstein generative adversarial network (VAW-GAN) that explicitly considers a VC objective when building the speech model. Experimental results corroborate the capability of our framework for building a VC system from unaligned data, and demonstrate improved conversion quality.

📄 PDF Abstract BibTeX arXiv:1704.00849

Code (1)

JeremyCCHsu/vae-npvc tf

Tasks

Generative Adversarial NetworkVoice Conversion

Similar Papers 제목 키워드 기반

Voice Conversion from Non-parallel Corpora Using Variational Auto-encoder

2016-10-13 · Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao 외

We propose a flexible framework for spectral conversion (SC) that facilitates training with unaligned corpora. Many SC frameworks require parallel corpora, phonetic alignments, or explicit frame-wise correspondence for l…

DecoderVoice Conversion

VAW-GAN for Singing Voice Conversion with Non-parallel Training Data

2020-08-10 · Junchen Lu, Kun Zhou, Berrak Sisman, Haizhou Li

Singing voice conversion aims to convert singer's voice from source to target without changing singing content. Parallel training data is typically required for the training of singing voice conversion system, that is ho…

DecoderGenerative Adversarial NetworkVoice Conversion

StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks

2018-06-06 · Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo

This paper proposes a method that allows non-parallel many-to-many voice conversion (VC) by using a variant of a generative adversarial network (GAN) called StarGAN. Our method, which we call StarGAN-VC, is noteworthy in…

AttributeGenerative Adversarial NetworkVoice Conversion

VAW-GAN for Disentanglement and Recomposition of Emotional Elements in Speech

2020-11-03 · Kun Zhou, Berrak Sisman, Haizhou Li

Emotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition…

DecoderDisentanglementGenerative Adversarial NetworkVoice Conversion

Semi-supervised voice conversion with amortized variational inference

2019-09-30 · Cory Stephenson, Gokce Keskin, Anil Thomas, Oguz H. Elibol

In this work we introduce a semi-supervised approach to the voice conversion problem, in which speech from a source speaker is converted into speech of a target speaker. The proposed method makes use of both parallel and…

Variational InferenceVoice Conversion