paper-with-me

Papers

wav2vec and its current potential to Automatic Speech Recognition in German for the usage in Digital History: A comparative assessment of available ASR-technologies for the use in cultural heritage contexts

2023-03-06 · Michael Fleck, Wolfgang Göderle

In this case study we trained and published a state-of-the-art open-source model for Automatic Speech Recognition (ASR) for German to evaluate the current potential of this technology for the use in the larger context of Digital Humanities and cultural heritage indexation. Along with this paper we publish our wav2vec2 based speech to text model while we evaluate its performance on a corpus of historical recordings we assembled compared against commercial cloud-based and proprietary services. While our model achieves moderate results, we see that proprietary cloud services fare significantly better. As our results show, recognition rates over 90 percent can currently be achieved, however, these numbers drop quickly once the recordings feature limited audio quality or use of non-every day or outworn language. A big issue is the high variety of different dialects and accents in the German language. Nevertheless, this paper highlights that the currently available quality of recognition is high enough to address various use cases in the Digital Humanities. We argue that ASR will become a key technology for the documentation and analysis of audio-visual sources and identify an array of important questions that the DH community and cultural heritage stakeholders will have to address in the near future.

📄 PDF Abstract BibTeX arXiv:2303.06026

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSpeech-to-Text

Similar Papers 제목 키워드 기반

LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition

2019-10-17 · LREC 2020 5 · Benjamin Beilharz, Xin Sun, Sariya Karimova, Stefan Riezler

We present a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks. The speech translation data consist of 110 hours of audio material aligned to over 50k pa…

Sentencespeech-recognitionSpeech RecognitionTranslation

Using Automatic Speech Recognition in Spoken Corpus Curation

2020-05-01 · LREC 2020 5 · Jan Gorisch, Michael Gref, Thomas Schmidt

The newest generation of speech technology caused a huge increase of audio-visual data nowadays being enhanced with orthographic transcripts such as in automatic subtitling in online platforms. Research data centers and …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation

2025-06-03 · Verena Blaschke, Miriam Winkler, Constantin Förster, Gabriele Wenger-Glemser 외

Although Germany has a diverse landscape of dialects, they are underrepresented in current automatic speech recognition (ASR) research. To enable studies of how robust models are towards dialectal variation, we present B…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Dialectal Speech Recognition and Translation of Swiss German Speech to Standard German Text: Microsoft's Submission to SwissText 2021

2021-06-15 · Yuriy Arabskyy, Aashish Agarwal, Subhadeep Dey, Oscar Koller

This paper describes the winning approach in the Shared Task 3 at SwissText 2021 on Swiss German Speech to Standard German Text, a public competition on dialect recognition and translation. Swiss German refers to the mul…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

German-Arabic Speech-to-Speech Translation for Psychiatric Diagnosis

2020-12-01 · COLING (WANLP) 2020 12 · Juan Hussain, Mohammed Mediani, Moritz Behr, M. Amin Cheragui 외

In this paper we present the natural language processing components of our German-Arabic speech-to-speech translation system which is being deployed in the context of interpretation during psychiatric, diagnostic intervi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDiagnostic+7