Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-Speech
The few-shot multi-speaker multi-style voice cloning task is to synthesize utterances with voice and speaking style similar to a reference speaker given only a few reference samples. In this work, we investigate different speaker representations and proposed to integrate pretrained and learnable speaker representations. Among different types of embeddings, the embedding pretrained by voice conversion achieves the best performance. The FastSpeech 2 model combined with both pretrained and learnable speaker representations shows great generalization ability on few-shot speakers and achieved 2nd place in the one-shot track of the ICASSP 2021 M2VoC challenge.
Code (1)
Tasks
text-to-speechText to SpeechVoice CloningVoice ConversionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation
This paper explores the use of ASR-pretrained Conformers for speaker verification, leveraging their strengths in modeling speech signals. We introduce three strategies: (1) Transfer learning to initialize the speaker emb…
Knowledge DistillationSpeaker VerificationTransfer LearningEmoji semantics/pragmatics: investigating commitment and lying
This paper presents the results of two experiments investigating the directness of emoji in constituting speaker meaning. This relationship is examined in two ways, with Experiment 1 testing whether speakers are committe…
FastAudio: A Learnable Audio Front-End for Spoof Speech Detection
Voice assistants, such as smart speakers, have exploded in popularity. It is currently estimated that the smart speaker adoption rate has exceeded 35% in the US adult population. Manufacturers have integrated speaker ide…
Speaker IdentificationSpeaker VerificationVoice Anti-spoofingInvestigating Speaker Embedding Disentanglement on Natural Read Speech
Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to increase the generalizability, explainabi…
DisentanglementFairnessRepresentation LearningSVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
Representations from pre-trained speech foundation models (SFMs) have shown impressive performance in many downstream tasks. However, the potential benefits of incorporating pre-trained SFM representations into speaker v…
Voice ConversionVoice Similarity