paper-with-me

홈 › Papers

gaHealth: An English–Irish Bilingual Corpus of Health Data

2022-06-01 · LREC 2022 6 · Séamus Lankford, Haithem Afli, Órla Ní Loinsigh, Andy Way

Machine Translation is a mature technology for many high-resource language pairs. However in the context of low-resource languages, there is a paucity of parallel data datasets available for developing translation models. Furthermore, the development of datasets for low-resource languages often focuses on simply creating the largest possible dataset for generic translation. The benefits and development of smaller in-domain datasets can easily be overlooked. To assess the merits of using in-domain data, a dataset for the specific domain of health was developed for the low-resource English to Irish language pair. Our study outlines the process used in developing the corpus and empirically demonstrates the benefits of using an in-domain dataset for the health domain. In the context of translating health-related data, models developed using the gaHealth corpus demonstrated a maximum BLEU score improvement of 22.2 points (40%) when compared with top performing models from the LoResMT2021 Shared Task. Furthermore, we define linguistic guidelines for developing gaHealth, the first bilingual corpus of health data for the Irish language, which we hope will be of use to other creators of low-resource data sets. gaHealth is now freely available online and is ready to be explored for further research.

📄 PDF Abstract BibTeX

Code (1)

seamusl/gahealth 공식 구현

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

gaHealth: An English-Irish Bilingual Corpus of Health Data

2024-03-06 · Séamus Lankford, Haithem Afli, Órla Ní Loinsigh, Andy Way

Machine Translation is a mature technology for many high-resource language pairs. However in the context of low-resource languages, there is a paucity of parallel data datasets available for developing translation models…

Machine TranslationTranslation

Enhancing Neural Machine Translation of Low-Resource Languages: Corpus Development, Human Evaluation and Explainable AI Architectures

2024-03-03 · Séamus Lankford

In the current machine translation (MT) landscape, the Transformer architecture stands out as the gold standard, especially for high-resource language pairs. This research delves into its efficacy for low-resource langua…

Machine TranslationTranslation

Qomhra: A Bilingual Irish and English Large Language Model

2025-10-20 · Joseph McInerney, Khanh-Tung Tran, Liam Lonergan, Ailbhe Ní Chasaide 외 arxiv

Large language model (LLM) research and development has overwhelmingly focused on the world's major languages, leading to under-representation of low-resource languages such as Irish. This paper introduces \textbf{Qomhrá…

Machine Translation in the Covid domain: an English-Irish case study for LoResMT 2021

2024-03-02 · MTSummit 2021 8 · Séamus Lankford, Haithem Afli, Andy Way

Translation models for the specific domain of translating Covid data from English to Irish were developed for the LoResMT 2021 shared task. Domain adaptation techniques, using a Covid-adapted generic 55k corpus from the …

8kDomain AdaptationMachine TranslationTranslation

Attentive fine-tuning of Transformers for Translation of low-resourced languages @LoResMT 2021

2021-08-19 · MTSummit 2021 8 · Karthik Puranik, Adeep Hande, Ruba Priyadharshini, Thenmozhi Durairaj 외

This paper reports the Machine Translation (MT) systems submitted by the IIITT team for the English->Marathi and English->Irish language pairs LoResMT 2021 shared task. The task focuses on getting exceptional translation…

Machine TranslationNMTTranslation