Versatile Speech Databases for High Quality Synthesis for Basque
This paper presents three new speech databases for standard Basque. They are designed primarily for corpus-based synthesis but each database has its specific purpose: 1) AhoSyn: high quality speech synthesis (recorded also in Spanish), 2) AhoSpeakers: voice conversion and 3) AhoEmo3: emotional speech synthesis. The whole corpus design and the recording process are described with detail. Once the databases were collected all the data was automatically labelled and annotated. Then, an HMM-based TTS voice was built and subjectively evaluated. The results of the evaluation are pretty satisfactory: 3.70 MOS for Basque and 3.44 for Spanish. Therefore, the evaluation assesses the quality of this new speech resource and the validity of the automated processing presented.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotional Speech SynthesisSpeech SynthesisVocal Bursts Intensity PredictionVoice ConversionSimilar Papers 제목 키워드 기반
Alert!... Calm Down, There is Nothing to Worry About. Warning and Soothing Speech Synthesis.
Presence of appropriate acoustic cues of affective features in the synthesized speech can be a prerequisite for the proper evaluation of the semantic content by the message recipient. In the recent work the authors have …
Expressive Speech SynthesisSentenceSpeech SynthesisDiffWave: A Versatile Diffusion Model for Audio Synthesis
In this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation. The model is non-autoregressive, and converts the white noise signal into structured wav…
Audio SynthesisDiversitymodelSpeech SynthesisEmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human …
text-to-speechText to SpeechA Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages
The performance of a text-to-speech (TTS) synthesis model depends on various factors, of which the quality of the training data is of utmost importance. Millions of data are collected around the globe for various languag…
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisHigh-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR
Automatic Speech Recognition (ASR) systems in the clinical domain face significant challenges, notably the need to recognise specialised medical vocabulary accurately and meet stringent precision requirements. We introdu…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1