Multi-speaker Emotional Text-to-speech Synthesizer
We present a methodology to train our multi-speaker emotional text-to-speech synthesizer that can express speech for 10 speakers' 7 different emotions. All silences from audio samples are removed prior to learning. This results in fast learning by our model. Curriculum learning is applied to train our model efficiently. Our model is first trained with a large single-speaker neutral dataset, and then trained with neutral speech from all speakers. Finally, our model is trained using datasets of emotional speech from all speakers. In each stage, training samples of each speaker-emotion pair have equal probability to appear in mini-batches. Through this procedure, our model can synthesize speech for all targeted speakers and emotions. Our synthesized audio sets are available on our web page.
Code (0)
등록된 구현이 없습니다.
Tasks
Alltext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source languages into English. We experiment wit…
Speech-to-Speech TranslationTranslationEmotional End-to-End Neural Speech Synthesizer
In this paper, we introduce an emotional speech synthesizer based on the recent end-to-end neural model, named Tacotron. Despite its benefits, we found that the original Tacotron suffers from the exposure bias problem an…
Zero-Shot Long-Form Voice Cloning with Dynamic Convolution Attention
With recent advancements in voice cloning, the performance of speech synthesis for a target speaker has been rendered similar to the human level. However, autoregressive voice cloning systems still suffer from text align…
FormSpeech Synthesistext-to-speechText to Speech+1Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS
Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the targe…
Language ModelingLanguage ModellingStyle Transfertext-to-speech+1ZET-Speech: Zero-shot adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models
Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional …
Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis