paper-with-me

홈 › Papers

RyanSpeech: A Corpus for Conversational Text-to-Speech Synthesis

2021-06-15 · Rohola Zandie, Mohammad H. Mahoor, Julia Madsen, Eshrat S. Emamian

This paper introduces RyanSpeech, a new speech corpus for research on automated text-to-speech (TTS) systems. Publicly available TTS corpora are often noisy, recorded with multiple speakers, or lack quality male speech data. In order to meet the need for a high quality, publicly available male speech corpus within the field of speech recognition, we have designed and created RyanSpeech which contains textual materials from real-world conversational settings. These materials contain over 10 hours of a professional male voice actor's speech recorded at 44.1 kHz. This corpus's design and pipeline make RyanSpeech ideal for developing TTS systems in real-world applications. To provide a baseline for future research, protocols, and benchmarks, we trained 4 state-of-the-art speech models and a vocoder on RyanSpeech. The results show 3.36 in mean opinion scores (MOS) in our best model. We have made both the corpus and trained models for public use.

📄 PDF Abstract BibTeX arXiv:2106.08468

Code (3)

roholazandie/ryan-tts 공식 구현 pytorch
Berthaniu/LatentOptimalPathsBayesianDP pytorch
xinleiniu/latentoptimalpathsbayesiandp pytorch

Tasks

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices

2022-06-01 · LREC 2022 6 · Liisi Piits, Hille Pajupuu, Heete Sahkai, Rene Altrov 외

Synthetic voices are increasingly used in applications that require a conversational speaking style, raising the question as to which type of training data yields the most suitable speaking style for such applications. T…

Sentencetext-to-speechText to Speech

DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech

2026-08-17 · Pengcheng Wang, Sheng Li, Jiyi Li, Takahiro Shinozaki arxiv

Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and t…

Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-based Multi-modal Context Modeling

2021-06-11 · Jingbei Li, Yi Meng, Chenyi Li, Zhiyong Wu 외

Comparing with traditional text-to-speech (TTS) systems, conversational TTS systems are required to synthesize speeches with proper speaking style confirming to the conversational context. However, state-of-the-art conte…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Evaluating expressive speech synthesis from audiobook corpora for conversational phrases

2012-05-01 · LREC 2012 5 · {\'E}va Sz{\'e}kely, Joao Paulo Cabral, Mohamed Abou-Zleikha, Peter Cahill 외

Audiobooks are a rich resource of large quantities of natural sounding, highly expressive speech. In our previous research we have shown that it is possible to detect different expressive voice styles represented in a pa…

ClusteringExpressive Speech SynthesisSpeech Synthesis

Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect

2026-06-24 · Dihia Lanasri, Rebeh Imane Ammar Aouchiche, Abdelkarim Remmide, Fairouz Taki 외 arxiv

Automatic speech and language technologies are still heavily biased toward high-resource languages, limiting their applicability to dialectal and low-resource settings such as Algerian Dialect. This language presents add…

Natural Language UnderstandingText-To-Speech SynthesisIntent ClassificationResponse Generation