paper-with-me

Papers

Many-to-Many Voice Conversion based Feature Disentanglement using Variational Autoencoder

2021-07-11 · Manh Luong, Viet Anh Tran

Voice conversion is a challenging task which transforms the voice characteristics of a source speaker to a target speaker without changing linguistic content. Recently, there have been many works on many-to-many Voice Conversion (VC) based on Variational Autoencoder (VAEs) achieving good results, however, these methods lack the ability to disentangle speaker identity and linguistic content to achieve good performance on unseen speaker scenarios. In this paper, we propose a new method based on feature disentanglement to tackle many to many voice conversion. The method has the capability to disentangle speaker identity and linguistic content from utterances, it can convert from many source speakers to many target speakers with a single autoencoder network. Moreover, it naturally deals with the unseen target speaker scenarios. We perform both objective and subjective evaluations to show the competitive performance of our proposed method compared with other state-of-the-art models in terms of naturalness and target speaker similarity.

📄 PDF Abstract BibTeX arXiv:2107.06642

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementVoice Conversion

Similar Papers 제목 키워드 기반

Phonetic Posteriorgrams based Many-to-Many Singing Voice Conversion via Adversarial Training

2020-12-03 · Haohan Guo, Heng Lu, Na Hu, Chunlei Zhang 외

This paper describes an end-to-end adversarial singing voice conversion (EA-SVC) approach. It can directly generate arbitrary singing waveform by given phonetic posteriorgram (PPG) representing content, F0 representing p…

Audio GenerationDisentanglementVoice Conversion

FragmentVC: Any-to-Any Voice Conversion by End-to-End Extracting and Fusing Fine-Grained Voice Fragments With Attention

2020-10-27 · Yist Y. Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung-Yi Lee 외

Any-to-any voice conversion aims to convert the voice from and to any speakers even unseen during training, which is much more challenging compared to one-to-one or many-to-many tasks, but much more attractive in real-wo…

DisentanglementSpeaker VerificationVoice Conversion

NVC-Net: End-to-End Adversarial Voice Conversion

2021-06-02 · Bac Nguyen, Fabien Cardinaux

Voice conversion has gained increasing popularity in many applications of speech synthesis. The idea is to change the voice identity from one speaker into another while keeping the linguistic content unchanged. Many voic…

GPUSpeech SynthesisVoice Conversion

Self-Supervised Representations for Singing Voice Conversion

2023-03-21 · Tejas Jayashankar, JiLong Wu, Leda Sari, David Kant 외

A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio representations such as HuBERT and Wav2Ve…

DisentanglementVoice Conversion

Many-to-Many Voice Conversion using Cycle-Consistent Variational Autoencoder with Multiple Decoders

2019-09-15 · Keonnyeong Lee, In-Chul Yoo, Dongsuk Yook

One of the obstacles in many-to-many voice conversion is the requirement of the parallel training data, which contain pairs of utterances with the same linguistic content spoken by different speakers. Since collecting su…

Voice Conversion