paper-with-me

Papers

Building a Morphological Network for Persian on Top of a Morpheme-Segmented Lexicon

2019-09-01 · WS 2019 9 · Hamid Haghdoost, Ebrahim Ansari, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}, Mahshid Nikravesh
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Supervised Morphological Segmentation Using Rich Annotated Lexicon

2019-09-01 · RANLP 2019 9 · Ebrahim Ansari, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}, Mohammad Mahmoudi, Hamid Haghdoost 외

Morphological segmentation of words is the process of dividing a word into smaller units called morphemes; it is tricky especially when a morphologically rich or polysynthetic language is under question. In this work, we…

Segmentation

CroDeriV: a new resource for processing Croatian morphology

2014-05-01 · LREC 2014 5 · Kre{\v{s}}imir {\v{S}}ojat, Matea Sreba{\v{c}}i{\'c}, Marko Tadi{\'c}, Tin Paveli{\'c}

The paper deals with the processing of Croatian morphology and presents CroDeriV ― a newly developed language resource that contains data about morphological structure and derivational relatedness of verbs in Croatian.…

LemmatizationMorphological Analysis

A Morphological Lexicon of Esperanto with Morpheme Frequencies

2016-05-01 · LREC 2016 5 · Eckhard Bick

This paper discusses the internal structure of complex Esperanto words (CWs). Using a morphological analyzer, possible affixation and compounding is checked for over 50,000 Esperanto lexemes against a list of 17,000 root…

POS

Automatic Detection of Morphological Processes in the Yorùbá Language

2022-06-01 · SIGUL (LREC) 2022 6 · Tunde Adegbola

Automatic morphology induction is important for computational processing of natural language. In resource-scarce languages in particular, it offers the possibility of supplementing data-driven strategies of Natural Langu…

SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers

2026-01-08 · Iaroslav Chelombitko, Ekaterina Chelombitko, Aleksey Komissarov arxiv

The quality of subword tokenization is critical for Large Language Models, yet evaluating tokenizers for morphologically rich Uralic languages is hampered by the lack of clean morpheme lexicons. We introduce SampoNLP, a …