A 500 Million Word POS-Tagged Icelandic Corpus
The new POS-tagged Icelandic corpus of the Leipzig Corpora Collection is an extensive resource for the analysis of the Icelandic language. As it contains a large share of all Web documents hosted under the .is top-level domain, it is especially valuable for investigations on modern Icelandic and non-standard language varieties. The corpus is accessible via a dedicated web portal and large shares are available for download. Focus of this paper will be the description of the tagging process and evaluation of statistical properties like word form frequencies and part of speech tag distributions. The latter will be in particular compared with values from the Icelandic Frequency Dictionary (IFD) Corpus.
Code (0)
등록된 구현이 없습니다.
Tasks
Part-Of-Speech TaggingPOSTAGSimilar Papers 제목 키워드 기반
Developing a Faroese PoS-tagging solution using Icelandic methods
We describe the development of a dedicated, high-accuracy part-of-speech (PoS) tagging solution for Faroese, a North Germanic language with about 50,000 speakers. To achieve this, a state-of-the-art neural PoS tagger for…
Part-Of-Speech TaggingPOSPOS TaggingCompiling and Filtering ParIce: An English-Icelandic Parallel Corpus
We present ParIce, a new English-Icelandic parallel corpus. This is the first parallel corpus built for the purposes of language technology development and research for Icelandic, although some Icelandic texts can be fou…
Nefnir: A high accuracy lemmatizer for Icelandic
Lemmatization, finding the basic morphological form of a word in a corpus, is an important step in many natural language processing tasks when working with morphologically rich languages. We describe and evaluate Nefnir,…
LemmatizationPOSVocal Bursts Intensity PredictionCorrecting Errors in a New Gold Standard for Tagging Icelandic Text
In this paper, we describe the correction of PoS tags in a new Icelandic corpus, MIM-GOLD, consisting of about 1 million tokens sampled from the Tagged Icelandic Corpus, M{\'I}M, released in 2013. The goal is to use the …
Part-Of-Speech TaggingPOSIGC-Parl: Icelandic Corpus of Parliamentary Proceedings
We describe the acquisition, annotation and encoding of the corpus of the Althingi parliamentary proceedings. The first version of the corpus includes speeches from 1911-2019. It comprises 406 thousand speeches and over …