paper-with-me

홈 › Papers

Qomhra: A Bilingual Irish and English Large Language Model

2025-10-20 · Joseph McInerney, Khanh-Tung Tran, Liam Lonergan, Ailbhe Ní Chasaide, Neasa Ní Chiaráin, Barry Devereux arxiv

Large language model (LLM) research and development has overwhelmingly focused on the world's major languages, leading to under-representation of low-resource languages such as Irish. This paper introduces \textbf{Qomhrá}, a bilingual Irish and English LLM, developed under extremely low-resource constraints. A complete pipeline is outlined spanning bilingual continued pre-training, instruction tuning, and the synthesis of human preference data for future alignment training. We focus on the lack of scalable methods to create human preference data by proposing a novel method to synthesise such data by prompting an LLM to generate `accepted'' and `rejected'' responses, which we validate as aligning with L1 Irish speakers. To select an LLM for synthesis, we evaluate the top closed-weight LLMs for Irish language generation performance. Gemini-2.5-Pro is ranked highest by L1 and L2 Irish-speakers, diverging from LLM-as-a-judge ratings, indicating a misalignment between current LLMs and the Irish-language community. Subsequently, we leverage Gemini-2.5-Pro to translate a large scale English-language instruction tuning dataset to Irish and to synthesise a first-of-its-kind Irish-language human preference dataset. We comprehensively evaluate Qomhrá across several benchmarks, testing translation, gender understanding, topic identification, and world knowledge; these evaluations show gains of up to 29\% in Irish and 44\% in English compared to the existing open-source Irish LLM baseline, UCCIX. The results of our framework provide insight and guidance to developing LLMs for both Irish and other low-resource languages.

📄 PDF Abstract BibTeX arXiv:2510.17652

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

gaHealth: An English–Irish Bilingual Corpus of Health Data

2022-06-01 · LREC 2022 6 · Séamus Lankford, Haithem Afli, Órla Ní Loinsigh, Andy Way

Machine Translation is a mature technology for many high-resource language pairs. However in the context of low-resource languages, there is a paucity of parallel data datasets available for developing translation models…

Machine TranslationTranslation

gaHealth: An English-Irish Bilingual Corpus of Health Data

2024-03-06 · Séamus Lankford, Haithem Afli, Órla Ní Loinsigh, Andy Way

Machine Translation is a mature technology for many high-resource language pairs. However in the context of low-resource languages, there is a paucity of parallel data datasets available for developing translation models…

Machine TranslationTranslation

AAC don Ghaeilge: the Prototype Development of Speech-Generating Assistive Technology for Irish

2022-06-01 · CLTW (LREC) 2022 6 · Emily Barnes, Oisín Morrin, Ailbhe Ní Chasaide, Julia Cummins 외

This paper describes the prototype development of an Alternative and Augmentative Communication (AAC) system for the Irish language. This system allows users to communicate using the ABAIR synthetic voices, by selecting …

Challenges in assistive technology development for an endangered language: an Irish (Gaelic) perspective

2022-05-01 · SLPAT (ACL) 2022 5 · Ailbhe Ni Chasaide, Emily Barnes, Neasa Ní Chiaráin, Ronan McGuirk 외

This paper describes three areas of assistive technology development which deploy the resources and speech technology for Irish (Gaelic), newly emerging from the ABAIR initiative. These include (i) a screenreading facili…

Attentive fine-tuning of Transformers for Translation of low-resourced languages @LoResMT 2021

2021-08-19 · MTSummit 2021 8 · Karthik Puranik, Adeep Hande, Ruba Priyadharshini, Thenmozhi Durairaj 외

This paper reports the Machine Translation (MT) systems submitted by the IIITT team for the English->Marathi and English->Irish language pairs LoResMT 2021 shared task. The task focuses on getting exceptional translation…

Machine TranslationNMTTranslation