paper-with-me

Papers

UCCIX: Irish-eXcellence Large Language Model

2024-05-13 · Khanh-Tung Tran, Barry O'Sullivan, Hoang D. Nguyen

The development of Large Language Models (LLMs) has predominantly focused on high-resource languages, leaving extremely low-resource languages like Irish with limited representation. This work presents UCCIX, a pioneering effort on the development of an open-source Irish-based LLM. We propose a novel framework for continued pre-training of LLMs specifically adapted for extremely low-resource languages, requiring only a fraction of the textual data typically needed for training LLMs according to scaling laws. Our model, based on Llama 2-13B, outperforms much larger models on Irish language tasks with up to 12% performance improvement, showcasing the effectiveness and efficiency of our approach. We also contribute comprehensive Irish benchmarking datasets, including IrishQA, a question-answering dataset, and Irish version of MT-bench. These datasets enable rigorous evaluation and facilitate future research in Irish LLM systems. Our work aims to preserve and promote the Irish language, knowledge, and culture of Ireland in the digital era while providing a framework for adapting LLMs to other indigenous languages.

📄 PDF Abstract BibTeX arXiv:2405.13010

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingLanguage ModelingLanguage ModellingLarge Language ModelmodelQuestion Answering

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

Qomhra: A Bilingual Irish and English Large Language Model

2025-10-20 · Joseph McInerney, Khanh-Tung Tran, Liam Lonergan, Ailbhe Ní Chasaide 외 arxiv

Large language model (LLM) research and development has overwhelmingly focused on the world's major languages, leading to under-representation of low-resource languages such as Irish. This paper introduces \textbf{Qomhrá…

Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting

2025-10-23 · Josh McGiff, Khanh-Tung Tran, William Mulcahy, Dáibhidh Ó Luinín 외 arxiv

We present Irish-BLiMP (Irish Benchmark of Linguistic Minimal Pairs), the first dataset and framework designed for fine-grained evaluation of linguistic competence in the Irish language, an endangered language. Drawing o…

TwittIrish: A Universal Dependencies Treebank of Tweets in Modern Irish

2022-05-01 · ACL 2022 5 · Lauren Cassidy, Teresa Lynn, James Barry, Jennifer Foster

Modern Irish is a minority language lacking sufficient computational resources for the task of accurate automatic syntactic parsing of user-generated content such as tweets. Although language technology for the Irish lan…

Dependency Parsing

TwittIrish: A Universal Dependencies Treebank of Tweets in Modern Irish

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Modern Irish is a minority language lacking sufficient linguistic resources for the task of accurate automatic syntactic parsing of user-generated content. As with other languages, the linguistic style observed in Irish…

IRIS: English-Irish Machine Translation System

2016-05-01 · LREC 2016 5 · Mihael Arcan, Caoilfhionn Lane, Eoin {\'O} Droighne{\'a}in, Paul Buitelaar

We describe IRIS, a statistical machine translation (SMT) system for translating from English into Irish and vice versa. Since Irish is considered an under-resourced language with a limited amount of machine-readable tex…

Machine TranslationTranslation