paper-with-me

Papers

The Norwegian Parliamentary Speech Corpus

2022-01-26 · LREC 2022 6 · Per Erik Solberg, Pablo Ortiz

The Norwegian Parliamentary Speech Corpus (NPSC) is a speech dataset with recordings of meetings from Stortinget, the Norwegian parliament. It is the first, publicly available dataset containing unscripted, Norwegian speech designed for training of automatic speech recognition (ASR) systems. The recordings are manually transcribed and annotated with language codes and speakers, and there are detailed metadata about the speakers. The transcriptions exist in both normalized and non-normalized form, and non-standardized words are explicitly marked and annotated with standardized equivalents. To test the usefulness of this dataset, we have compared an ASR system trained on the NPSC with a baseline system trained on only manuscript-read speech. These systems were tested on an independent dataset containing spontaneous, dialectal speech. The NPSC-trained system performed significantly better, with a 22.9% relative improvement in word error rate (WER). Moreover, training on the NPSC is shown to have a "democratizing" effect in terms of dialects, as improvements are generally larger for dialects with higher WER from the baseline system.

📄 PDF Abstract BibTeX arXiv:2201.10881

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Boosting Norwegian Automatic Speech Recognition

2023-07-04 · Javier de la Rosa, Rolv-Arild Braaten, Per Egil Kummervold, Freddy Wetjen 외

In this paper, we present several baselines for automatic speech recognition (ASR) models for the two official written languages in Norway: Bokm{\aa}l and Nynorsk. We compare the performance of models of varying sizes an…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Models

2025-09-23 · Erik Božík, Marek Šuppa arxiv

Slovak remains a low-resource language for automatic speech recognition (ASR), with fewer than 100 hours of publicly available training data. We present SloPal, a comprehensive Slovak parliamentary corpus comprising 330,…

Speech Recognition

An open source part-of-speech tagger for Norwegian: Building on existing language resources

2014-05-01 · LREC 2014 5 · Cristina S{\'a}nchez Marco

This paper presents an open source part-of-speech tagger for the Norwegian language. It describes how an existing language processing library (FreeLing) was used to build a new part-of-speech tagger for this language. Th…

Dependency ParsingMachine TranslationMorphological AnalysisMorphological Tagging+1

The Corpus of American Norwegian Speech (CANS)

2015-05-01 · WS 2015 5 · Janne Bondi Johannessen

IGC-Parl: Icelandic Corpus of Parliamentary Proceedings

2020-05-01 · LREC 2020 5 · Stein{\th}{\'o}r Steingr{\'\i}msson, Starka{\dh}ur Barkarson, Gunnar Thor {\"O}rn{\'o}lfsson

We describe the acquisition, annotation and encoding of the corpus of the Althingi parliamentary proceedings. The first version of the corpus includes speeches from 1911-2019. It comprises 406 thousand speeches and over …