paper-with-me

Papers

Versatile Speech Databases for High Quality Synthesis for Basque

2012-05-01 · LREC 2012 5 · I{\~n}aki Sainz, Daniel Erro, Eva Navas, Inma Hern{\'a}ez, Jon Sanchez, Ibon Saratxaga, Igor Odriozola

This paper presents three new speech databases for standard Basque. They are designed primarily for corpus-based synthesis but each database has its specific purpose: 1) AhoSyn: high quality speech synthesis (recorded also in Spanish), 2) AhoSpeakers: voice conversion and 3) AhoEmo3: emotional speech synthesis. The whole corpus design and the recording process are described with detail. Once the databases were collected all the data was automatically labelled and annotated. Then, an HMM-based TTS voice was built and subjectively evaluated. The results of the evaluation are pretty satisfactory: 3.70 MOS for Basque and 3.44 for Spanish. Therefore, the evaluation assesses the quality of this new speech resource and the validity of the automated processing presented.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Emotional Speech SynthesisSpeech SynthesisVocal Bursts Intensity PredictionVoice Conversion

Similar Papers 제목 키워드 기반

Alert!... Calm Down, There is Nothing to Worry About. Warning and Soothing Speech Synthesis.

2014-05-01 · LREC 2014 5 · Milan Rusko, Sakhia Darjaa, Mari{\'a}n Trnka, Mari{\'a}n Ritomsk{\'y} 외

Presence of appropriate acoustic cues of affective features in the synthesized speech can be a prerequisite for the proper evaluation of the semantic content by the message recipient. In the recent work the authors have …

Expressive Speech SynthesisSentenceSpeech Synthesis

DiffWave: A Versatile Diffusion Model for Audio Synthesis

2020-09-21 · ICLR 2021 1 · Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 외

In this work, we propose DiffWave, a versatile diffusion probabilistic model for conditional and unconditional waveform generation. The model is non-autoregressive, and converts the white noise signal into structured wav…

Audio SynthesisDiversitymodelSpeech Synthesis

EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations

2024-12-09 · Weizhen Bian, Yubo Zhou, Kaitai Zhang, Xiaohan Gu

Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human …

text-to-speechText to Speech

A Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages

2024-10-18 · Sujitha Sathiyamoorthy, N Mohana, Anusha Prakash, Hema A Murthy

The performance of a text-to-speech (TTS) synthesis model depends on various factors, of which the quality of the training data is of utmost importance. Millions of data are collected around the globe for various languag…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR

2024-11-24 · Sourav Banerjee, Ayushi Agarwal, Promila Ghosh

Automatic Speech Recognition (ASR) systems in the clinical domain face significant challenges, notably the need to recognise specialised medical vocabulary accurately and meet stringent precision requirements. We introdu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1