paper-with-me

Papers

Open Source Automatic Speech Recognition for German

2018-07-26 · Benjamin Milde, Arne Köhn

High quality Automatic Speech Recognition (ASR) is a prerequisite for speech-based applications and research. While state-of-the-art ASR software is freely available, the language dependent acoustic models are lacking for languages other than English, due to the limited amount of freely available training data. We train acoustic models for German with Kaldi on two datasets, which are both distributed under a Creative Commons license. The resulting model is freely redistributable, lowering the cost of entry for German ASR. The models are trained on a total of 412 hours of German read speech data and we achieve a relative word error reduction of 26% by adding data from the Spoken Wikipedia Corpus to the previously best freely available German acoustic model recipe and dataset. Our best model achieves a word error rate of 14.38 on the Tuda-De test set. Due to the large amount of speakers and the diversity of topics included in the training data, our model is robust against speaker variation and topic shift.

📄 PDF Abstract BibTeX arXiv:1807.10311

Code (2)

uhh-lt/kaldi-tuda-de 공식 구현
tudarmstadt-lt/kaldi-tuda-de

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Open Source German Distant Speech Recognition: Corpus and Acoustic Model

2015-12-11 · International Conference on Text, Speech, and Dialogue 2015 12 · Stephan Radeck-Arneth, Benjamin Milde, Arvid Lange, Evandro Gouvea 외

We present a new freely available corpus for German distant speech recognition and report speaker-independent word error rate (WER) results for two open source speech recognizers trained on this corpus. The corpus has be…

Distant Speech Recognitionspeech-recognitionSpeech Recognition

LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition

2019-10-17 · LREC 2020 5 · Benjamin Beilharz, Xin Sun, Sariya Karimova, Stefan Riezler

We present a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks. The speech translation data consist of 110 hours of audio material aligned to over 50k pa…

Sentencespeech-recognitionSpeech RecognitionTranslation

wav2vec and its current potential to Automatic Speech Recognition in German for the usage in Digital History: A comparative assessment of available ASR-technologies for the use in cultural heritage contexts

2023-03-06 · Michael Fleck, Wolfgang Göderle

In this case study we trained and published a state-of-the-art open-source model for Automatic Speech Recognition (ASR) for German to evaluate the current potential of this technology for the use in the larger context of…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Effects of Layer Freezing on Transferring a Speech Recognition System to Under-resourced Languages

2021-02-08 · KONVENS (WS) 2021 9 · Onno Eberhard, Torsten Zesch

In this paper, we investigate the effect of layer freezing on the effectiveness of model transfer in the area of automatic speech recognition. We experiment with Mozilla's DeepSpeech architecture on German and Swiss Germ…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Exploiting the large-scale German Broadcast Corpus to boost the Fraunhofer IAIS Speech Recognition System

2014-05-01 · LREC 2014 5 · Michael Stadtschnitzer, Jochen Schwenninger, Daniel Stein, Joachim Koehler

In this paper we describe the large-scale German broadcast corpus (GER-TV1000h) containing more than 1,000 hours of transcribed speech data. This corpus is unique in the German language corpora domain and enables signifi…

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Modelling+5