UQuAD1.0: Development of an Urdu Question Answering Training Data for Machine Reading Comprehension
In recent years, low-resource Machine Reading Comprehension (MRC) has made significant progress, with models getting remarkable performance on various language datasets. However, none of these models have been customized for the Urdu language. This work explores the semi-automated creation of the Urdu Question Answering Dataset (UQuAD1.0) by combining machine-translated SQuAD with human-generated samples derived from Wikipedia articles and Urdu RC worksheets from Cambridge O-level books. UQuAD1.0 is a large-scale Urdu dataset intended for extractive machine reading comprehension tasks consisting of 49k question Answers pairs in question, passage, and answer format. In UQuAD1.0, 45000 pairs of QA were generated by machine translation of the original SQuAD1.0 and approximately 4000 pairs via crowdsourcing. In this study, we used two types of MRC models: rule-based baseline and advanced Transformer-based models. However, we have discovered that the latter outperforms the others; thus, we have decided to concentrate solely on Transformer-based architectures. Using XLMRoBERTa and multi-lingual BERT, we acquire an F1 score of 0.66 and 0.63, respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesMachine Reading ComprehensionMachine TranslationQuestion AnsweringTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
The development of Urdu scene text detection, recognition, and Visual Question Answering (VQA) technologies is crucial for advancing accessibility, information retrieval, and linguistic diversity in digital content, faci…
DiversityInformation RetrievalQuestion AnsweringRetrieval+4UQA: Corpus for Urdu Question Answering
This paper introduces UQA, a novel dataset for question answering and text comprehension in Urdu, a low-resource language with over 70 million native speakers. UQA is generated by translating the Stanford Question Answer…
Multilingual NLPQuestion AnsweringReading ComprehensionMultilingual Hematology Visual Question Answering Dataset
Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for tasks such as Visual Question Answering. However, existing hematology …
Visual Question AnsweringLEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
We present LEGAL-UQA, the first Urdu legal question-answering dataset derived from Pakistan's constitution. This parallel English-Urdu dataset includes 619 question-answer pairs, each with corresponding legal article con…
Optical Character Recognition (OCR)Question AnsweringRetrievalUrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
The rapid use of large language models (LLMs) has raised critical concerns regarding the factual reliability of their outputs, especially in low-resource languages such as Urdu. Existing automated fact-checking solutions…
BenchmarkingClaim VerificationFact CheckingQuestion Answering+1