Creating Data in Icelandic for Text Normalization
There is no natural way to acquire normalized data so we try to create good enough data to attempt more advanced methods for text normalization. We manually annotated the first normalized corpus in Icelandic, 40,000 sentences, and developed Regína, a rule-based system for text normalization. Regína gets 90.83% accuracy compared to the manually annotated corpus on non-standard words. Regína showed a significant improvement in accuracy when compared to an older normalization system for Icelandic. The normalized corpus and Regína will be released as open source.
Code (0)
등록된 구현이 없습니다.
Tasks
Text NormalizationSimilar Papers 제목 키워드 기반
Creating a Parallel Icelandic Dependency Treebank from Raw Text to Universal Dependencies
Making the low-resource language, Icelandic, accessible and usable in Language Technology is a work in progress and is supported by the Icelandic government. Creating resources and suitable training data (e.g., a depende…
Natural Questions in Icelandic
We present the first extractive question answering (QA) dataset for Icelandic, Natural Questions in Icelandic (NQiI). Developing such datasets is important for the development and evaluation of Icelandic QA systems. It a…
Extractive Question-AnsweringNatural QuestionsQuestion AnsweringCompiling and Filtering ParIce: An English-Icelandic Parallel Corpus
We present ParIce, a new English-Icelandic parallel corpus. This is the first parallel corpus built for the purposes of language technology development and research for Icelandic, although some Icelandic texts can be fou…
An Evaluation of Neural Machine Translation Models on Historical Spelling Normalization
In this paper, we apply different NMT models to the problem of historical spelling normalization for five languages: English, German, Hungarian, Icelandic, and Swedish. The NMT models are at different levels, have differ…
Machine TranslationNMTTranslationA Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models
We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error…
Constituency ParsingGrammatical Error Detectionnamed-entity-recognitionNamed Entity Recognition+3