Adapitch: Adaption Multi-Speaker Text-to-Speech Conditioned on Pitch Disentangling with Untranscribed Data
In this paper, we proposed Adapitch, a multi-speaker TTS method that makes adaptation of the supervised module with untranscribed data. We design two self supervised modules to train the text encoder and mel decoder separately with untranscribed data to enhance the representation of text and mel. To better handle the prosody information in a synthesized voice, a supervised TTS module is designed conditioned on content disentangling of pitch, text, and speaker. The training phase was separated into two parts, pretrained and fixed the text encoder and mel decoder with unsupervised mode, then the supervised mode on the disentanglement of TTS. Experiment results show that the Adaptich achieved much better quality than baseline methods.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderDisentanglementtext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis
In this paper, we propose a method of speaker adaption with intuitive prosodic features for statistical parametric speech synthesis. The intuitive prosodic features employed in this method include pitch, pitch range, spe…
Speech SynthesisCUHK-EE Voice Cloning System for ICASSP 2021 M2VoC Challenge
This paper presents the CUHK-EE voice cloning system for ICASSP 2021 M2VoC challenge. The challenge provides two Mandarin speech corpora: the AIShell-3 corpus of 218 speakers with noise and reverberation and the MST corp…
Voice CloningFactorised Speaker-environment Adaptive Training of Conformer Speech Recognition Systems
Rich sources of variability in natural speech present significant challenges to current data intensive speech recognition technologies. To model both speaker and environment level diversity, this paper proposes a novel B…
Diversityspeech-recognitionSpeech RecognitionTest-time AdaptationU-vectors: Generating clusterable speaker embedding from unlabeled data
Speaker recognition deals with recognizing speakers by their speech. Most speaker recognition systems are built upon two stages, the first stage extracts low dimensional correlation embeddings from speech, and the second…
Domain AdaptationSpeaker RecognitionThe 2016 KIT IWSLT Speech-to-Text Systems for English and German
This paper describes our German and English Speech-to-Text (STT) systems for the 2016 IWSLT evaluation campaign. The campaign focuses on the transcription of unsegmented TED talks. Our setup includes systems using both t…
Speech-to-Text