Neural Speech Synthesis in German
While many speech synthesis systems based on deep neural networks are thoroughly evaluated and released for free use in English, models for languages with far less active speakers like German are scarcely trained and most often not published for common use. This work covers specific challenges in training text to speech models for the German language, including dataset selection and data preprocessing, and presents the training process for multiple models of an end-to-end text to speech system based on a combination of Tacotron 2 and Multi-Band MelGAN. All model compositions were evaluated against the mean opinion score, which revealed comparable results to models in literature that are trained and evaluated on English datasets. In addition, empirical analyses identified distinct aspects influencing the quality of such systems, based on subjective user experience. All trained models are released for public use.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisSimilar Papers 제목 키워드 기반
Text-to-Speech Pipeline for Swiss German -- A comparison
In this work, we studied the synthesis of Swiss German speech using different Text-to-Speech (TTS) models. We evaluated the TTS models on three corpora, and we found, that VITS models performed best, hence, using them fo…
Speech Synthesistext-to-speechText to SpeechGerman-Arabic Speech-to-Speech Translation for Psychiatric Diagnosis
In this paper we present the natural language processing components of our German-Arabic speech-to-speech translation system which is being deployed in the context of interpretation during psychiatric, diagnostic intervi…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDiagnostic+7SDS-200: A Swiss German Speech to Standard German Text Corpus
We present SDS-200, a corpus of Swiss German dialectal speech with Standard German text translations, annotated with dialect, age, and gender information of the speakers. The dataset allows for training speech translatio…
Speech SynthesisTranslationSprachsynthese -- State-of-the-Art in englischer und deutscher Sprache
Reading text aloud is an important feature for modern computer applications. It not only facilitates access to information for visually impaired people, but is also a pleasant convenience for non-impaired users. In this …
Speech SynthesisBuilding Text-to-Speech Systems for Resource Poor Languages
This paper describes research on building text-to-speech synthesis systems (TTS) for resource poor languages using available resources from other languages and describes our general approach to building cross-linguistic …
ClusteringSpeech Synthesistext-to-speechText to Speech+1