paper-with-me

홈 › Papers

Samrómur Children: An Icelandic Speech Corpus

2022-06-01 · LREC 2022 6 · Carlos Daniel Hernandez Mena, David Erik Mollberg, Michal Borský, Jón Guðnason

Samrómur Children is an Icelandic speech corpus intended for the field of automatic speech recognition. It contains 131 hours of read speech from Icelandic children aged between 4 to 17 years. The test portion was meticulously selected to cover a wide range of ages as possible; we aimed to have exactly the same amount of data per age range. The speech was collected with the crowd-sourcing platform Samrómur.is, which is inspired on the “Mozilla’s Common Voice Project”. The corpus was developed within the framework of the “Language Technology Programme for Icelandic 2019 − 2023”; the goal of the project is to make Icelandic available in language-technology applications. Samrómur Children is the first corpus in Icelandic with children’s voices for public use under a Creative Commons license. Additionally, we present baseline experiments and results using Kaldi.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Samr\'omur: Crowd-sourcing Data Collection for Icelandic Speech Recognition

2020-05-01 · LREC 2020 5 · David Erik Mollberg, {\'O}lafur Helgi J{\'o}nsson, Sunneva {\TH}orsteinsd{\'o}ttir, Stein{\th}{\'o}r Steingr{\'\i}msson 외

This contribution describes an ongoing project of speech data collection, using the web application Samr{\'o}mur which is built upon Common Voice, Mozilla Foundation{'}s web platform for open-source voice collection. The…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Marketingspeech-recognition+1

A 500 Million Word POS-Tagged Icelandic Corpus

2014-05-01 · LREC 2014 5 · Thomas Eckart, Erla Hallsteinsd{\'o}ttir, Sigr{\'u}n Helgad{\'o}ttir, Uwe Quasthoff 외

The new POS-tagged Icelandic corpus of the Leipzig Corpora Collection is an extensive resource for the analysis of the Icelandic language. As it contains a large share of all Web documents hosted under the .is top-level …

Part-Of-Speech TaggingPOSTAG

Parsing Icelandic Al\thingi Transcripts: Parliamentary Speeches as a Genre

2020-05-01 · LREC 2020 5 · Kristj{\'a}n R{\'u}narsson, Einar Freyr Sigur{\dh}sson

We introduce a corpus of transcripts from Al{\th}ingi, the Icelandic parliament. The corpus is syntactically parsed for phrase structure according to the annotation scheme of the Icelandic Parsed Historical Corpus (IcePa…

A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models

2022-01-14 · Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir 외

We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error…

Constituency ParsingGrammatical Error Detectionnamed-entity-recognitionNamed Entity Recognition+3

A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models

2022-06-01 · LREC 2022 6 · Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir 외

We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error…

Constituency ParsingGrammatical Error Detectionnamed-entity-recognitionNamed Entity Recognition+3