paper-with-me

Papers

Szeged Corpus 2.5: Morphological Modifications in a Manually POS-tagged Hungarian Corpus

2014-05-01 · LREC 2014 5 · Veronika Vincze, Viktor Varga, Katalin Ilona Simk{\'o}, J{\'a}nos Zsibrita, {\'A}goston Nagy, Rich{\'a}rd Farkas, J{\'a}nos Csirik

The Szeged Corpus is the largest manually annotated database containing the possible morphological analyses and lemmas for each word form. In this work, we present its latest version, Szeged Corpus 2.5, in which the new harmonized morphological coding system of Hungarian has been employed and, on the other hand, the majority of misspelled words have been corrected and tagged with the proper morphological code. New morphological codes are introduced for participles, causative / modal / frequentative verbs, adverbial pronouns and punctuation marks, moreover, the distinction between common and proper nouns is eliminated. We also report some statistical data on the frequency of the new morphological codes. The new version of the corpus made it possible to train magyarlanc, a data-driven POS-tagger of Hungarian on a dataset with the new harmonized codes. According to the results, magyarlanc is able to achieve a state-of-the-art accuracy score on the 2.5 version as well.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

POS

Similar Papers 제목 키워드 기반

Morphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models

2019-12-02 · Nelda Kote, Marenglen Biba, Jenna Kanerva, Samuel Rönnqvist 외

In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…

LemmatizationMorphological TaggingPart-Of-Speech Tagging

Universal Dependencies and Morphology for Hungarian - and on the Price of Universality

2017-04-01 · EACL 2017 4 · Veronika Vincze, Katalin Simk{\'o}, Zsolt Sz{\'a}nt{\'o}, Rich{\'a}rd Farkas

In this paper, we present how the principles of universal dependencies and morphology have been adapted to Hungarian. We report the most challenging grammatical phenomena and our solutions to those. On the basis of the a…

Morphological Tagging

Creating a morphological and syntactic tagged corpus for the Uzbek language

2022-10-27 · Maksud Sharipov, Jamolbek Mattiev, Jasur Sobirov, Rustam Baltayev

Nowadays, creation of the tagged corpora is becoming one of the most important tasks of Natural Language Processing (NLP). There are not enough tagged corpora to build machine learning models for the low-resource Uzbek l…

POS

Nefnir: A high accuracy lemmatizer for Icelandic

2019-07-27 · WS (NoDaLiDa) 2019 9 · Svanhvít Lilja Ingólfsdóttir, Hrafn Loftsson, Jón Friðrik Daðason, Kristín Bjarnadóttir

Lemmatization, finding the basic morphological form of a word in a corpus, is an important step in many natural language processing tasks when working with morphologically rich languages. We describe and evaluate Nefnir,…

LemmatizationPOSVocal Bursts Intensity Prediction

TALC-sef A Manually-Revised POS-TAgged Literary Corpus in Serbian, English and French

2014-05-01 · LREC 2014 5 · Antonio Balvet, Dejan Stosic, Aleks Miletic, ra

In this paper, we present a parallel literary corpus for Serbian, English and French, the TALC-sef corpus. The corpus includes a manually-revised pos-tagged reference Serbian corpus of over 150,000 words. The initial obj…

POSPOS Tagging