Disentangled Speech Representation Learning for One-Shot Cross-lingual Voice Conversion Using $β$-VAE
We propose an unsupervised learning method to disentangle speech into content representation and speaker identity representation. We apply this method to the challenging one-shot cross-lingual voice conversion task to demonstrate the effectiveness of the disentanglement. Inspired by $\beta$-VAE, we introduce a learning objective that balances between the information captured by the content and speaker representations. In addition, the inductive biases from the architectural design and the training dataset further encourage the desired disentanglement. Both objective and subjective evaluations show the effectiveness of the proposed method in speech disentanglement and in one-shot cross-lingual voice conversion.
Code (0)
등록된 구현이 없습니다.
Tasks
DisentanglementRepresentation LearningSpeech Representation LearningVoice ConversionSimilar Papers 제목 키워드 기반
Investigation of using disentangled and interpretable representations with language conditioning for cross-lingual voice conversion
We study the problem of cross-lingual voice conversion in non-parallel speech corpora and one-shot learning setting. Most prior work require either parallel speech corpora or enough amount of training data from a target …
One-Shot LearningVoice ConversionCrossSpeech: Speaker-independent Acoustic Representation for Cross-lingual Speech Synthesis
While recent text-to-speech (TTS) systems have made remarkable strides toward human-level quality, the performance of cross-lingual TTS lags behind that of intra-lingual TTS. This gap is mainly rooted from the speaker-la…
Speech Synthesistext-to-speechText to SpeechParrotTTS: Text-to-Speech synthesis by exploiting self-supervised representations
We present ParrotTTS, a modularized text-to-speech synthesis model leveraging disentangled self-supervised speech representations. It can train a multi-speaker variant effectively using transcripts from a single speaker.…
Self-Supervised LearningSpeech Synthesistext-to-speechText to Speech+1Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
We propose a novel description-based controllable text-to-speech (TTS) method with cross-lingual control capability. To address the lack of audio-description paired data in the target language, we combine a TTS model tra…
Self-Supervised Learningtext-to-speechText to SpeechT-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation
We present a new approach to perform zero-shot cross-modal transfer between speech and text for translation tasks. Multilingual speech and text are encoded in a joint fixed-size representation space. Then, we compare dif…
DecoderMachine Translationtext-to-speechText to Speech+2