paper-with-me

홈 › Papers

Creating Data in Icelandic for Text Normalization

2021-05-01 · NoDaLiDa 2021 5 · Helga Svala Sigurðardóttir, Anna Björk Nikulásdóttir, Jón Guðnason

There is no natural way to acquire normalized data so we try to create good enough data to attempt more advanced methods for text normalization. We manually annotated the first normalized corpus in Icelandic, 40,000 sentences, and developed Regína, a rule-based system for text normalization. Regína gets 90.83% accuracy compared to the manually annotated corpus on non-standard words. Regína showed a significant improvement in accuracy when compared to an older normalization system for Icelandic. The normalized corpus and Regína will be released as open source.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Text Normalization

Similar Papers 제목 키워드 기반

Creating a Parallel Icelandic Dependency Treebank from Raw Text to Universal Dependencies

2020-05-01 · LREC 2020 5 · Hildur J{\'o}nsd{\'o}ttir, Anton Karl Ingason

Making the low-resource language, Icelandic, accessible and usable in Language Technology is a work in progress and is supported by the Icelandic government. Creating resources and suitable training data (e.g., a depende…

Natural Questions in Icelandic

2022-06-01 · LREC 2022 6 · Vésteinn Snæbjarnarson, Hafsteinn Einarsson

We present the first extractive question answering (QA) dataset for Icelandic, Natural Questions in Icelandic (NQiI). Developing such datasets is important for the development and evaluation of Icelandic QA systems. It a…

Extractive Question-AnsweringNatural QuestionsQuestion Answering

Compiling and Filtering ParIce: An English-Icelandic Parallel Corpus

2019-09-01 · WS (NoDaLiDa) 2019 9 · Starkaður Barkarson, Steinþór Steingrímsson

We present ParIce, a new English-Icelandic parallel corpus. This is the first parallel corpus built for the purposes of language technology development and research for Icelandic, although some Icelandic texts can be fou…

An Evaluation of Neural Machine Translation Models on Historical Spelling Normalization

2018-06-13 · COLING 2018 8 · Gongbo Tang, Fabienne Cap, Eva Pettersson, Joakim Nivre

In this paper, we apply different NMT models to the problem of historical spelling normalization for five languages: English, German, Hungarian, Icelandic, and Swedish. The NMT models are at different levels, have differ…

Machine TranslationNMTTranslation

A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models

2022-01-14 · Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir 외

We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error…

Constituency ParsingGrammatical Error Detectionnamed-entity-recognitionNamed Entity Recognition+3