Ossetic-COT: Designing a morphologically annotated corpus and morphological analyzer for Ossetic
In this work we present the first morphologically annotated corpus for Iron Ossetic that conforms to the Universal Dependencies schema. The corpus includes 5454 manually annotated sentences from the Iron Ossetic Corpus of Oral Texts, containing 74032 tokens. We use this corpus to train a BERT-based morphological analyzer. The analyzer achieves tag accuracy of 95.60%.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A Morphologically Annotated Hebrew CHILDES Corpus
A Morphologically Annotated Corpus of Emirati Arabic
Neural disambiguation of lemma and part of speech in morphologically rich languages
We consider the problem of disambiguating the lemma and part of speech of ambiguous words in morphologically rich languages. We propose a method for disambiguating ambiguous words in context, using a large un-annotated c…
LEMMAPOSSupervised Morphological Segmentation Using Rich Annotated Lexicon
Morphological segmentation of words is the process of dividing a word into smaller units called morphemes; it is tricky especially when a morphologically rich or polysynthetic language is under question. In this work, we…
SegmentationMorphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models
In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…
LemmatizationMorphological TaggingPart-Of-Speech Tagging