paper-with-me

Papers

GLeMM: A large-scale multilingual dataset for morphological research

2026-04-14 · Hathout Nabil, Basilio Calderone, Fiammetta Namer, Franck Sajous arxiv

In derivational morphology, what mechanisms govern the variation in form-meaning relations between words? The answers to this type of questions are typically based on intuition and on observations drawn from limited data, even when a wide range of languages is considered. Many of these studies are difficult to replicate and generalize. To address this issue, we present GLeMM, a new derivational resource designed for experimentation and data-driven description in morphology. GLeMM is characterized by (i) its large size, (ii) its extensive coverage (currently amounting to seven European languages, i.e., German, English, Spanish, French, Italian, Polish, Russian, (iii) its fully automated design, identical across all languages, (iv) the automatic annotation of morphological features on each entry, as well as (v) the encoding of semantic descriptions for a significant subset of these entries. It enables researchers to address difficult questions, such as the role of form and meaning in word-formation, and to develop and experimentally test computational methods that identify the structures of derivational morphology. The article describes how GLeMM is created using Wiktionary articles and presents various case studies illustrating possible applications of the resource.

📄 PDF Abstract BibTeX arXiv:2604.12442

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation

2026-04-20 · Mehul Agarwal, Aditya Aggarwal, Arnav Goel, Medha Hira 외 arxiv

While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their ability to handle grammatical gender and morphological agreement remains underexplored. In m…

Question Answering

MorphyNet: a Large Multilingual Database of Derivational and Inflectional Morphology

2021-08-01 · ACL (SIGMORPHON) 2021 8 · Khuyagbaatar Batsuren, Gábor Bella, Fausto Giunchiglia

Large-scale morphological databases provide essential input to a wide range of NLP applications. Inflectional data is of particular importance for morphologically rich (agglutinative and highly inflecting) languages, and…

YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models

2026-09-09 · Mahmoud Reda, Salam Khalifa, Reham Marzouk, Nizar Habash arxiv

Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly…

Evaluation of the Accuracy of the BGLemmatizer

2015-06-13 · Elena Karashtranova, Grigor Iliev, Nadezhda Borisova, Yana Chankova 외

This paper reveals the results of an analysis of the accuracy of developed software for automatic lemmatization for the Bulgarian language. This lemmatization software is written entirely in Java and is distributed as a …

Lemmatization

Morphemes Without Borders: Evaluating Root-Pattern Morphology in Arabic Tokenizers and LLMs

2026-03-16 · Yara Alakeel, Chatrine Qwaider, Hanan Aldarmaki, Sawsan Alqahtani arxiv

This work investigates how effectively large language models (LLMs) and their tokenization schemes represent and generate Arabic root-pattern morphology, probing whether they capture genuine morphological structure or re…