Speaker Generation
This work explores the task of synthesizing speech in nonexistent human-sounding voices. We call this task "speaker generation", and present TacoSpawn, a system that performs competitively at this task. TacoSpawn is a recurrent attention-based text-to-speech model that learns a distribution over a speaker embedding space, which enables sampling of novel and diverse speakers. Our method is easy to implement, and does not require transfer learning from speaker ID systems. We present objective and subjective metrics for evaluating performance on this task, and demonstrate that our proposed objective metrics correlate with human perception of speaker similarity. Audio samples are available on our demo page.
Code (0)
등록된 구현이 없습니다.
Tasks
text-to-speechText to SpeechTransfer LearningSimilar Papers 제목 키워드 기반
Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation
Recent advancements in personalized speech generation have brought synthetic speech increasingly close to the realism of target speakers' recordings, yet multimodal speaker generation remains on the rise. This paper intr…
DiversityWe Need Variations in Speech Generation: Sub-center Modelling for Speaker Embeddings
Modeling the rich prosodic variations inherent in human speech is essential for generating natural-sounding speech. While speaker embeddings are commonly used as conditioning inputs in personalized speech generation, the…
Speaker RecognitionSpeech SynthesisVoice ConversionImproving the generation of personalised descriptions
Referring expression generation (REG) models that use speaker-dependent information require a considerable amount of training data produced by every individual speaker, or may otherwise perform poorly. In this work we pr…
Referring ExpressionReferring expression generationText GenerationAdversarial speech for voice privacy protection from Personalized Speech generation
The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human list…
Speaker Verificationtext-to-speechText to SpeechVoice ConversionCommunity Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization
The clustering algorithm plays a crucial role in speaker diarization systems. However, traditional clustering algorithms suffer from the complex distribution of speaker embeddings and lack of digging potential relationsh…
ClusteringCommunity DetectionGraph Generationspeaker-diarization+1