paper-with-me

홈 › Papers

Adapitch: Adaption Multi-Speaker Text-to-Speech Conditioned on Pitch Disentangling with Untranscribed Data

2022-10-25 · xulong Zhang, Jianzong Wang, Ning Cheng, Jing Xiao

In this paper, we proposed Adapitch, a multi-speaker TTS method that makes adaptation of the supervised module with untranscribed data. We design two self supervised modules to train the text encoder and mel decoder separately with untranscribed data to enhance the representation of text and mel. To better handle the prosody information in a synthesized voice, a supervised TTS module is designed conditioned on content disentangling of pitch, text, and speaker. The training phase was separated into two parts, pretrained and fixed the text encoder and mel decoder with unsupervised mode, then the supervised mode on the disentanglement of TTS. Experiment results show that the Adaptich achieved much better quality than baseline methods.

📄 PDF Abstract BibTeX arXiv:2210.13803

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDisentanglementtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis

2022-03-02 · Pengyu Cheng, ZhenHua Ling

In this paper, we propose a method of speaker adaption with intuitive prosodic features for statistical parametric speech synthesis. The intuitive prosodic features employed in this method include pitch, pitch range, spe…

Speech Synthesis

CUHK-EE Voice Cloning System for ICASSP 2021 M2VoC Challenge

2021-03-08 · Daxin Tan, Hingpang Huang, Guangyan Zhang, Tan Lee

This paper presents the CUHK-EE voice cloning system for ICASSP 2021 M2VoC challenge. The challenge provides two Mandarin speech corpora: the AIShell-3 corpus of 218 speakers with noise and reverberation and the MST corp…

Voice Cloning

Factorised Speaker-environment Adaptive Training of Conformer Speech Recognition Systems

2023-06-26 · Jiajun Deng, Guinan Li, Xurong Xie, Zengrui Jin 외

Rich sources of variability in natural speech present significant challenges to current data intensive speech recognition technologies. To model both speaker and environment level diversity, this paper proposes a novel B…

Diversityspeech-recognitionSpeech RecognitionTest-time Adaptation

U-vectors: Generating clusterable speaker embedding from unlabeled data

2021-02-07 · M. F. Mridha, Abu Quwsar Ohi, Muhammad Mostafa Monowar, Md. Abdul Hamid 외

Speaker recognition deals with recognizing speakers by their speech. Most speaker recognition systems are built upon two stages, the first stage extracts low dimensional correlation embeddings from speech, and the second…

Domain AdaptationSpeaker Recognition

The 2016 KIT IWSLT Speech-to-Text Systems for English and German

2016-12-01 · IWSLT 2016 12 · Thai-Son Nguyen, Markus Müller, Matthias Sperber, Thomas Zenkel 외

This paper describes our German and English Speech-to-Text (STT) systems for the 2016 IWSLT evaluation campaign. The campaign focuses on the transcription of unsegmented TED talks. Our setup includes systems using both t…

Speech-to-Text