paper-with-me

홈 › Papers

RUSLAN: Russian Spoken Language Corpus for Speech Synthesis

2019-06-26 · Lenar Gabdrakhmanov, Rustem Garaev, Evgenii Razinkov

We present RUSLAN -- a new open Russian spoken language corpus for the text-to-speech task. RUSLAN contains 22200 audio samples with text annotations -- more than 31 hours of high-quality speech of one person -- being the largest annotated Russian corpus in terms of speech duration for a single speaker. We trained an end-to-end neural network for the text-to-speech task on our corpus and evaluated the quality of the synthesized speech using Mean Opinion Score test. Synthesized speech achieves 4.05 score for naturalness and 3.78 score for intelligibility on a 5-point MOS scale.

📄 PDF Abstract BibTeX arXiv:1906.11645

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

TheRuSLan: Database of Russian Sign Language

2020-05-01 · LREC 2020 5 · Ildar Kagirov, Denis Ivanko, Dmitry Ryumin, Alex Axyonov 외

In this paper, a new Russian sign language multimedia database TheRuSLan is presented. The database includes lexical units (single words and phrases) from Russian sign language within one subject area, namely, {``}food p…

Gesture RecognitionSign Language Recognition

MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible

2019-07-30 · LREC 2020 5 · Marcely Zanon Boito, William N. Havard, Mahault Garnerin, Éric Le Ferrand 외

The CMU Wilderness Multilingual Speech Dataset (Black, 2019) is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)RetrievalSentence+5

Minimally Supervised Written-to-Spoken Text Normalization

2016-09-21 · Ke Wu, Kyle Gorman, Richard Sproat

In speech-applications such as text-to-speech (TTS) or automatic speech recognition (ASR), \emph{text normalization} refers to the task of converting from a \emph{written} representation into a representation of how the …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+3

CoRuSS - a New Prosodically Annotated Corpus of Russian Spontaneous Speech

2016-05-01 · LREC 2016 5 · Tatiana Kachkovskaia, Daniil Kocharov, Pavel Skrelin, Nina Volskaya

This paper describes speech data recording, processing and annotation of a new speech corpus CoRuSS (Corpus of Russian Spontaneous Speech), which is based on connected communicative speech recorded from 60 native Russian…

Numbers Normalisation in the Inflected Languages: a Case Study of Polish

2019-08-01 · WS 2019 8 · Rafa{\l} Po{\'s}wiata, Micha{\l} Pere{\l}kiewicz

Text normalisation in Text-to-Speech systems is a process of converting written expressions to their spoken forms. This task is complicated because in many cases the normalised form depends on the context. Furthermore, w…

text-to-speechText to Speech