Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
This paper focuses on adapting the functionalities of the FastPitch model to the Romanian language; extending the set of speakers from one to eighteen; synthesising speech using an anonymous identity; and replicating the identities of new, unseen speakers. During this work, the effects of various configurations and training strategies were tested and discussed, along with their advantages and weaknesses. Finally, we settled on a new configuration, built on top of the FastPitch architecture, capable of producing natural speech synthesis, for both known (identities from the training dataset) and unknown (identities learnt through short reference samples) speakers. The anonymous speaker can be used for text-to-speech synthesis, if one wants to cancel out the identity information while keeping the semantic content whole and clear. At last, we discussed possible limitations of our work, which will form the basis for future investigations and advancements.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Auto Spell Suggestion for High Quality Speech Synthesis in Hindi
The goal of Text-to-Speech (TTS) synthesis in a particular language is to convert arbitrary input text to intelligible and natural sounding speech. However, for a particular language like Hindi, which is a highly confusi…
Speech Synthesistext-to-speechText to SpeechVocal Bursts Intensity PredictionLoRP-TTS: Low-Rank Personalized Text-To-Speech
Speech synthesis models convert written text into natural-sounding audio. While earlier models were limited to a single speaker, recent advancements have led to the development of zero-shot systems that generate realisti…
Speech Synthesistext-to-speechText to SpeechTowards Controllable Speech Synthesis in the Era of Large Language Models: A Survey
Text-to-speech (TTS), also known as speech synthesis, is a prominent research area that aims to generate natural-sounding human speech from text. Recently, with the increasing industrial demand, TTS technologies have evo…
Speech SynthesisSurveytext-to-speechText to SpeechSpeech Recognition with Augmented Synthesized Speech
Recent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribe…
Data AugmentationDiversityRobust Speech Recognitionspeech-recognition+2An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis
For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning. There are more and more deep learning-based TTS systems developed to make it possible to produce voices with …
Speech Synthesis