paper-with-me

홈 › Papers

A Corpus of Read and Spontaneous Upper Saxon German Speech for ASR Evaluation

2016-05-01 · LREC 2016 5 · Robert Herms, Laura Seelig, Stefanie M{\"u}nch, Maximilian Eibl

In this Paper we present a corpus named SXUCorpus which contains read and spontaneous speech of the Upper Saxon German dialect. The data has been collected from eight archives of local television stations located in the Free State of Saxony. The recordings include broadcasted topics of news, economy, weather, sport, and documentation from the years 1992 to 1996 and have been manually transcribed and labeled. In the paper, we report the methodology of collecting and processing analog audiovisual material, constructing the corpus and describe the properties of the data. In its current version, the corpus is available to the scientific community and is designed for automatic speech recognition (ASR) evaluation with a development set and a test set. We performed ASR experiments with the open-source framework sphinx-4 including a configuration for Standard German on the dataset. Additionally, we show the influence of acoustic model and language model adaptation by the utilization of the development set.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Low Saxon dialect distances at the orthographic and syntactic level

2022-05-01 · LChange (ACL) 2022 5 · Janine Siewert, Yves Scherrer, Martijn Wieling

We compare five Low Saxon dialects from the 19th and 21st century from Germany and the Netherlands with each other as well as with modern Standard Dutch and Standard German. Our comparison is based on character n-grams o…

POS

SwissGPC v1.0 -- The Swiss German Podcasts Corpus

2025-09-24 · Samuel Stucki, Mark Cieliebak, Jan Deriu arxiv

We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to ta…

GRASS: the Graz corpus of Read And Spontaneous Speech

2014-05-01 · LREC 2014 5 · Barbara Schuppler, Martin Hagmueller, Juan A. Morales-Cordovilla, Hannes Pessentheiner

This paper provides a description of the preparation, the speakers, the recordings, and the creation of the orthographic transcriptions of the first large scale speech database for Austrian German. It contains approximat…

Speech Recognition

FOLK-Gold ― A Gold Standard for Part-of-Speech-Tagging of Spoken German

2016-05-01 · LREC 2016 5 · Swantje Westpfahl, Thomas Schmidt

In this paper, we present a GOLD standard of part-of-speech tagged transcripts of spoken German. The GOLD standard data consists of four annotation layers ― transcription (modified orthography), normalization (standard o…

LemmatizationPart-Of-Speech TaggingPOS

LSDC - A comprehensive dataset for Low Saxon Dialect Classification

2020-12-01 · VarDial (COLING) 2020 12 · Janine Siewert, Yves Scherrer, Martijn Wieling, Jörg Tiedemann

We present a new comprehensive dataset for the unstandardised West-Germanic language Low Saxon covering the last two centuries, the majority of modern dialects and various genres, which will be made openly available in c…

Classification