paper-with-me

홈 › Papers

Irish National Morphology Database: a high-accuracy open-source dataset of Irish words

2014-08-01 · WS 2014 8 · Michal Boleslav M{\v{e}}chura
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MoirfEolas and CríochScore: Developing Resources for and the Evaluation of Tokenization Alignment with Irish Morphology

2026-09-04 · Jane Adkins, Abigail Walsh, Brian Davis, Elaine Uí Dhonnchadha arxiv

This paper presents new tokenization resources for Irish and evaluation measures of alignment with the morphological boundaries of the language. We present MoirfEolas, a dataset of over 35,000 Irish words mapped to their…

Introducing the National Corpus of Irish Project

2022-06-01 · CLTW (LREC) 2022 6 · Mícheál Ó Meachair, Úna Bhreathnach, Gearóid Ó Cleircín

This paper introduces the National Corpus of Irish, an initiative to develop a large national corpus of written and spoken contemporary Irish as well as related specialised corpora. The newly-compiled corpora will be hos…

LEMMA

Developing high-end reusable tools and resources for Irish-language terminology, lexicography, onomastics (toponymy), folkloristics, and more, using modern web and database technologies

2014-08-01 · WS 2014 8 · Brian {\'O} Raghallaigh, Michal Boleslav M{\v{e}}chura

Bayesian Causal Forests for Multivariate Outcomes: Application to Irish Data From an International Large Scale Education Assessment

2023-03-08 · Nathan McJames, Andrew Parnell, Yong Chen Goh, Ann O'Shea

Bayesian Causal Forests (BCF) is a causal inference machine learning model based on a highly flexible non-parametric regression and classification tool called Bayesian Additive Regression Trees (BART). Motivated by data …

Causal Inferenceregression

Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting

2025-10-23 · Josh McGiff, Khanh-Tung Tran, William Mulcahy, Dáibhidh Ó Luinín 외 arxiv

We present Irish-BLiMP (Irish Benchmark of Linguistic Minimal Pairs), the first dataset and framework designed for fine-grained evaluation of linguistic competence in the Irish language, an endangered language. Drawing o…