EuSQuAD: Automatically Translated and Aligned SQuAD2.0 for Basque
The widespread availability of Question Answering (QA) datasets in English has greatly facilitated the advancement of the Natural Language Processing (NLP) field. However, the scarcity of such resources for minority languages, such as Basque, poses a substantial challenge for these communities. In this context, the translation and alignment of existing QA datasets plays a crucial role in narrowing this technological gap. This work presents EuSQuAD, the first initiative dedicated to automatically translating and aligning SQuAD2.0 into Basque, resulting in more than 142k QA examples. We demonstrate EuSQuAD's value through extensive qualitative analysis and QA experiments supported with EuSQuAD as training data. These experiments are evaluated with a new human-annotated dataset.
Code (1)
Tasks
Question AnsweringSimilar Papers 제목 키워드 기반
Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on Basque
Large language models (LLMs) are typically optimized for resource-rich languages like English, exacerbating the gap between high-resource and underrepresented languages. This work presents a detailed analysis of strategi…
Instruction FollowingNatural Language UnderstandingAdding the Basque Parliament Corpus to ParlaMint Project
The aim of this work is to describe the colection created with transcript of the Basque parliamentary speeches. This corpus follows the constraints of the ParlaMint project. The Basque ParlaMint corpus consists of two ve…
The ADAPT System Description for the IWSLT 2018 Basque to English Translation Task
In this paper we present the ADAPT system built for the Basque to English Low Resource MT Evaluation Campaign. Basque is a low-resourced, morphologically-rich language. This poses a challenge for Neural Machine Translati…
Machine TranslationTranslationReading Comprehension in Czech via Machine Translation and Cross-lingual Transfer
Reading comprehension is a well studied task, with huge training datasets in English. This work focuses on building reading comprehension systems for Czech, without requiring any manually annotated Czech training data. F…
Cross-Lingual TransferMachine TranslationReading ComprehensionTranslationAmaSQuAD: A Benchmark for Amharic Extractive Question Answering
This research presents a novel framework for translating extractive question-answering datasets into low-resource languages, as demonstrated by the creation of the AmaSQuAD dataset, a translation of SQuAD 2.0 into Amhari…
Extractive Question-AnsweringQuestion AnsweringXLM-R