DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority group listeners, and useful across various applications and context. Speech synthesis can further be made more flexible by allowing users to choose any combination of speaker identity and accent, resulting in a wide range of personalized speech outputs. Current models struggle to disentangle speaker and accent representation, making it difficult to accurately imitate different accents while maintaining the same speaker characteristics. We propose a novel approach to disentangle speaker and accent representations using multi-level variational autoencoders (ML-VAE) and vector quantization (VQ) to improve flexibility and enhance personalization in speech synthesis. Our proposed method addresses the challenge of effectively separating speaker and accent characteristics, enabling more fine-grained control over the synthesized speech. Code and speech samples are publicly available.
Code (0)
등록된 구현이 없습니다.
Tasks
DisentanglementQuantizationSpeech Synthesistext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Multilingual Multiaccented Multispeaker TTS with RADTTS
We work to create a multilingual speech synthesis system which can generate speech with the proper accent while retaining the characteristics of an individual voice. This is challenging to do because it is expensive to o…
Speech SynthesisZero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network
The goal of accent conversion (AC) is to convert the accent of speech into the target accent while preserving the content and speaker identity. AC enables a variety of applications, such as language learning, speech cont…
Data AugmentationDisentanglementCross-lingual Multispeaker Text-to-Speech under Limited-Data Scenario
Modeling voices for multiple speakers and multiple languages in one text-to-speech system has been a challenge for a long time. This paper presents an extension on Tacotron2 to achieve bilingual multispeaker speech synth…
AttributeSpeech Synthesistext-to-speechText to SpeechSpeaker verification-derived loss and data augmentation for DNN-based multispeaker speech synthesis
Building multispeaker neural network-based text-to-speech synthesis systems commonly relies on the availability of large amounts of high quality recordings from each speaker and conditioning the training process on the s…
Data AugmentationSpeaker VerificationSpeech Synthesistext-to-speech+2Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
Generating speech across different accents while preserving speaker identity is crucial for various real-world applications. However, accurately and independently modeling both speaker and accent characteristics in text-…
DisentanglementSpeech Synthesistext-to-speechText to Speech+1