paper-with-me

홈 › Papers

A 500 Million Word POS-Tagged Icelandic Corpus

2014-05-01 · LREC 2014 5 · Thomas Eckart, Erla Hallsteinsd{\'o}ttir, Sigr{\'u}n Helgad{\'o}ttir, Uwe Quasthoff, Dirk Goldhahn

The new POS-tagged Icelandic corpus of the Leipzig Corpora Collection is an extensive resource for the analysis of the Icelandic language. As it contains a large share of all Web documents hosted under the .is top-level domain, it is especially valuable for investigations on modern Icelandic and non-standard language varieties. The corpus is accessible via a dedicated web portal and large shares are available for download. Focus of this paper will be the description of the tagging process and evaluation of statistical properties like word form frequencies and part of speech tag distributions. The latter will be in particular compared with values from the Icelandic Frequency Dictionary (IFD) Corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech TaggingPOSTAG

Similar Papers 제목 키워드 기반

Developing a Faroese PoS-tagging solution using Icelandic methods

2020-12-01 · ICON 2020 12 · Hinrik Hafsteinsson, Anton Karl Ingason

We describe the development of a dedicated, high-accuracy part-of-speech (PoS) tagging solution for Faroese, a North Germanic language with about 50,000 speakers. To achieve this, a state-of-the-art neural PoS tagger for…

Part-Of-Speech TaggingPOSPOS Tagging

Compiling and Filtering ParIce: An English-Icelandic Parallel Corpus

2019-09-01 · WS (NoDaLiDa) 2019 9 · Starkaður Barkarson, Steinþór Steingrímsson

We present ParIce, a new English-Icelandic parallel corpus. This is the first parallel corpus built for the purposes of language technology development and research for Icelandic, although some Icelandic texts can be fou…

Nefnir: A high accuracy lemmatizer for Icelandic

2019-07-27 · WS (NoDaLiDa) 2019 9 · Svanhvít Lilja Ingólfsdóttir, Hrafn Loftsson, Jón Friðrik Daðason, Kristín Bjarnadóttir

Lemmatization, finding the basic morphological form of a word in a corpus, is an important step in many natural language processing tasks when working with morphologically rich languages. We describe and evaluate Nefnir,…

LemmatizationPOSVocal Bursts Intensity Prediction

Correcting Errors in a New Gold Standard for Tagging Icelandic Text

2014-05-01 · LREC 2014 5 · Sigr{\'u}n Helgad{\'o}ttir, Hrafn Loftsson, Eir{\'\i}kur R{\"o}gnvaldsson

In this paper, we describe the correction of PoS tags in a new Icelandic corpus, MIM-GOLD, consisting of about 1 million tokens sampled from the Tagged Icelandic Corpus, M{\'I}M, released in 2013. The goal is to use the …

Part-Of-Speech TaggingPOS

IGC-Parl: Icelandic Corpus of Parliamentary Proceedings

2020-05-01 · LREC 2020 5 · Stein{\th}{\'o}r Steingr{\'\i}msson, Starka{\dh}ur Barkarson, Gunnar Thor {\"O}rn{\'o}lfsson

We describe the acquisition, annotation and encoding of the corpus of the Althingi parliamentary proceedings. The first version of the corpus includes speeches from 1911-2019. It comprises 406 thousand speeches and over …