paper-with-me

Papers

A Morphological Lexicon of Esperanto with Morpheme Frequencies

2016-05-01 · LREC 2016 5 · Eckhard Bick

This paper discusses the internal structure of complex Esperanto words (CWs). Using a morphological analyzer, possible affixation and compounding is checked for over 50,000 Esperanto lexemes against a list of 17,000 root words. Morpheme boundaries in the resulting analyses were then checked manually, creating a CW dictionary of 28,000 words, representing 56.4{\%} of the lexicon, or 19.4{\%} of corpus tokens. The error percentage of the EspGram morphological analyzer for new corpus CWs was 4.3{\%} for types and 6.4{\%} for tokens, with a recall of almost 100{\%}, and wrong/spurious boundaries being more common than missing ones. For pedagogical purposes a morpheme frequency dictionary was constructed for a 16 million word corpus, confirming the importance of agglutinative derivational morphemes in the Esperanto lexicon. Finally, as a means to reduce the morphological ambiguity of CWs, we provide POS likelihoods for Esperanto suffixes.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

POS

Similar Papers 제목 키워드 기반

CroDeriV: a new resource for processing Croatian morphology

2014-05-01 · LREC 2014 5 · Kre{\v{s}}imir {\v{S}}ojat, Matea Sreba{\v{c}}i{\'c}, Marko Tadi{\'c}, Tin Paveli{\'c}

The paper deals with the processing of Croatian morphology and presents CroDeriV ― a newly developed language resource that contains data about morphological structure and derivational relatedness of verbs in Croatian.…

LemmatizationMorphological Analysis

Building a Morphological Network for Persian on Top of a Morpheme-Segmented Lexicon

2019-09-01 · WS 2019 9 · Hamid Haghdoost, Ebrahim Ansari, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}, Mahshid Nikravesh

Supervised Morphological Segmentation Using Rich Annotated Lexicon

2019-09-01 · RANLP 2019 9 · Ebrahim Ansari, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}, Mohammad Mahmoudi, Hamid Haghdoost 외

Morphological segmentation of words is the process of dividing a word into smaller units called morphemes; it is tricky especially when a morphologically rich or polysynthetic language is under question. In this work, we…

Segmentation

Automatic Detection of Morphological Processes in the Yorùbá Language

2022-06-01 · SIGUL (LREC) 2022 6 · Tunde Adegbola

Automatic morphology induction is important for computational processing of natural language. In resource-scarce languages in particular, it offers the possibility of supplementing data-driven strategies of Natural Langu…

SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers

2026-01-08 · Iaroslav Chelombitko, Ekaterina Chelombitko, Aleksey Komissarov arxiv

The quality of subword tokenization is critical for Large Language Models, yet evaluating tokenizers for morphologically rich Uralic languages is hampered by the lack of clean morpheme lexicons. We introduce SampoNLP, a …