paper-with-me

홈 › Papers

Ossetic-COT: Designing a morphologically annotated corpus and morphological analyzer for Ossetic

2026-07-06 · Anna Shatskikh, Alexey Sorokin arxiv

In this work we present the first morphologically annotated corpus for Iron Ossetic that conforms to the Universal Dependencies schema. The corpus includes 5454 manually annotated sentences from the Iron Ossetic Corpus of Oral Texts, containing 74032 tokens. We use this corpus to train a BERT-based morphological analyzer. The analyzer achieves tag accuracy of 95.60%.

📄 PDF Abstract BibTeX arXiv:2607.04895

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Morphologically Annotated Hebrew CHILDES Corpus

2012-04-01 · WS 2012 4 · Aviad Albert, Brian MacWhinney, Bracha Nir, Shuly Wintner
Language Acquisition

A Morphologically Annotated Corpus of Emirati Arabic

2018-05-01 · LREC 2018 5 · Salam Khalifa, Nizar Habash, Fadhl Eryani, Ossama Obeid 외
LemmatizationMachine TranslationMorphological AnalysisPart-Of-Speech Tagging

Neural disambiguation of lemma and part of speech in morphologically rich languages

2020-07-12 · LREC 2020 5 · José María Hoya Quecedo, Maximilian W. Koppatz, Giacomo Furlan, Roman Yangarber

We consider the problem of disambiguating the lemma and part of speech of ambiguous words in morphologically rich languages. We propose a method for disambiguating ambiguous words in context, using a large un-annotated c…

LEMMAPOS

Supervised Morphological Segmentation Using Rich Annotated Lexicon

2019-09-01 · RANLP 2019 9 · Ebrahim Ansari, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}, Mohammad Mahmoudi, Hamid Haghdoost 외

Morphological segmentation of words is the process of dividing a word into smaller units called morphemes; it is tricky especially when a morphologically rich or polysynthetic language is under question. In this work, we…

Segmentation

Morphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models

2019-12-02 · Nelda Kote, Marenglen Biba, Jenna Kanerva, Samuel Rönnqvist 외

In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…

LemmatizationMorphological TaggingPart-Of-Speech Tagging