paper-with-me

홈 › Papers

Designing the Latvian Speech Recognition Corpus

2014-05-01 · LREC 2014 5 · M{\=a}rcis Pinnis, Ilze Auzi{\c{n}}a, K{\=a}rlis Goba

In this paper the authors present the first Latvian speech corpus designed specifically for speech recognition purposes. The paper outlines the decisions made in the corpus designing process through analysis of related work on speech corpora creation for different languages. The authors provide also guidelines that were used for the creation of the Latvian speech recognition corpus. The corpus creation guidelines are fairly general for them to be re-used by other researchers when working on different language speech recognition corpora. The corpus consists of two parts ― an orthographically annotated corpus containing 100 hours of orthographically transcribed audio data and a phonetically annotated corpus containing 4 hours of phonetically transcribed audio data. Metadata files in XML format provide additional details about the speakers, noise levels, speech styles, etc. The speech recognition corpus is phonetically balanced and phonetically rich and the paper describes also the methodology how the phonetical balancedness has been assessed.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionSpeech Synthesis

Similar Papers 제목 키워드 기반

Designing a Speech Corpus for the Development and Evaluation of Dictation Systems in Latvian

2016-05-01 · LREC 2016 5 · M{\=a}rcis Pinnis, Askars Salimbajevs, Ilze Auzi{\c{n}}a

In this paper the authors present a speech corpus designed and created for the development and evaluation of dictation systems in Latvian. The corpus consists of over nine hours of orthographically annotated speech from …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+3

Development and Evaluation of Speech Synthesis Corpora for Latvian

2020-05-01 · LREC 2020 5 · Roberts Dar{\c{g}}is, Peteris Paikens, Normunds Gruzitis, Ilze Auzina 외

Text to speech (TTS) systems are necessary for all languages to ensure accessibility and availability of digital language services. Recent advances in neural speech synthesis have eText to speech (TTS) systems are necess…

speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speech+1

Error Analysis and Improving Speech Recognition for Latvian Language

2015-09-01 · RANLP 2015 9 · Askars Salimbajevs, Jevgenijs Strigins
speech-recognitionSpeech Recognition

Using sub-word n-gram models for dealing with OOV in large vocabulary speech recognition for Latvian

2015-05-01 · WS 2015 5 · Askars Salimbajevs, Jevgenijs Strigins
Language Modellingspeech-recognitionSpeech Recognition

LaVA – Latvian Language Learner corpus

2022-06-01 · LREC 2022 6 · Roberts Darģis, Ilze Auziņa, Inga Kaija, Kristīne Levāne-Petrova 외

This paper presents the Latvian Language Learner Corpus (LaVA) developed at the Institute of Mathematics and Computer Science, University of Latvia. LaVA corpus contains 1015 essays (190k tokens and 790k characters exclu…