Samrómur Children: An Icelandic Speech Corpus
Samrómur Children is an Icelandic speech corpus intended for the field of automatic speech recognition. It contains 131 hours of read speech from Icelandic children aged between 4 to 17 years. The test portion was meticulously selected to cover a wide range of ages as possible; we aimed to have exactly the same amount of data per age range. The speech was collected with the crowd-sourcing platform Samrómur.is, which is inspired on the “Mozilla’s Common Voice Project”. The corpus was developed within the framework of the “Language Technology Programme for Icelandic 2019 − 2023”; the goal of the project is to make Icelandic available in language-technology applications. Samrómur Children is the first corpus in Icelandic with children’s voices for public use under a Creative Commons license. Additionally, we present baseline experiments and results using Kaldi.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Samr\'omur: Crowd-sourcing Data Collection for Icelandic Speech Recognition
This contribution describes an ongoing project of speech data collection, using the web application Samr{\'o}mur which is built upon Common Voice, Mozilla Foundation{'}s web platform for open-source voice collection. The…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Marketingspeech-recognition+1A 500 Million Word POS-Tagged Icelandic Corpus
The new POS-tagged Icelandic corpus of the Leipzig Corpora Collection is an extensive resource for the analysis of the Icelandic language. As it contains a large share of all Web documents hosted under the .is top-level …
Part-Of-Speech TaggingPOSTAGParsing Icelandic Al\thingi Transcripts: Parliamentary Speeches as a Genre
We introduce a corpus of transcripts from Al{\th}ingi, the Icelandic parliament. The corpus is syntactically parsed for phrase structure according to the annotation scheme of the Icelandic Parsed Historical Corpus (IcePa…
A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models
We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error…
Constituency ParsingGrammatical Error Detectionnamed-entity-recognitionNamed Entity Recognition+3A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models
We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error…
Constituency ParsingGrammatical Error Detectionnamed-entity-recognitionNamed Entity Recognition+3