paper-with-me

홈 › Papers

MoirfEolas and CríochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology

2026-09-04 · Jane Adkins, Abigail Walsh, Brian Davis, Elaine Uí Dhonnchadha arxiv

This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphological boundaries of the language. We present MoirfEolas, a dataset of over 35,000 Irish words mapped to their respective eclipses, prefixes and suffixes as well as an evaluation metric CríochScore, that evaluates the alignment of tokenizations with the morphological boundaries present in MoirfEolas. We evaluate common tokenization algorithms using CríochScore as well as intrinsic metrics present in the tokenization literature. We find that the Unigram Language Model aligns with Irish morphology more often than the other algorithms evaluated. We also find trade-offs between morphological-alignment of tokenization with both compression as well as vocabulary efficiency, providing practical insights for Irish natural language processing development. This dataset contributes towards combating the Irish language's low-resource status; moreover, the construction process reported in this paper can be emulated by other languages to create specialised morphological resources.

📄 PDF Abstract BibTeX arXiv:2609.05022

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Foundations and Evaluations in NLP

2025-04-02 · Jungyeul Park

This memoir explores two fundamental aspects of Natural Language Processing (NLP): the creation of linguistic resources and the evaluation of NLP system performance. Over the past decade, my work has focused on developin…

Boundary DetectionDependency Parsingnamed-entity-recognitionNamed Entity Recognition+2

BPEmb: Tokenization-free Pre-trained Subword Embeddings in 275 Languages

2017-10-05 · LREC 2018 5 · Benjamin Heinzerling, Michael Strube

We present BPEmb, a collection of pre-trained subword unit embeddings in 275 languages, based on Byte-Pair Encoding (BPE). In an evaluation using fine-grained entity typing as testbed, BPEmb performs competitively, and f…

Entity TypingWord Embeddings

Vacillating Human Correlation of SacreBLEU in Unprotected Languages

2022-05-01 · HumEval (ACL) 2022 5 · Ahrii Kim, Jinhyeon Kim

SacreBLEU, by incorporating a text normalizing step in the pipeline, has become a rising automatic evaluation metric in recent MT studies. With agglutinative languages such as Korean, however, the lexical-level metric ca…

Contextual morphologically-guided tokenization for Latin encoder models

2025-11-12 · Marisa Hudspeth, Patrick J. Burns, Brendan O'Connor arxiv

Tokenization is a critical component of language model pretraining, yet standard tokenization methods often prioritize information-theoretical goals like high compression and low fertility rather than linguistic goals li…

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

2025-02-10 · M. Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Sercan Karakaş 외

Tokenization is a fundamental preprocessing step in NLP, directly impacting large language models' (LLMs) ability to capture syntactic, morphosyntactic, and semantic structures. This paper introduces a novel framework fo…

MMLUMorphological AnalysisMultiple-choicevalid