paper-with-me

홈 › Papers

UQuAD1.0: Development of an Urdu Question Answering Training Data for Machine Reading Comprehension

2021-11-02 · Samreen Kazi, Shakeel Khoja

In recent years, low-resource Machine Reading Comprehension (MRC) has made significant progress, with models getting remarkable performance on various language datasets. However, none of these models have been customized for the Urdu language. This work explores the semi-automated creation of the Urdu Question Answering Dataset (UQuAD1.0) by combining machine-translated SQuAD with human-generated samples derived from Wikipedia articles and Urdu RC worksheets from Cambridge O-level books. UQuAD1.0 is a large-scale Urdu dataset intended for extractive machine reading comprehension tasks consisting of 49k question Answers pairs in question, passage, and answer format. In UQuAD1.0, 45000 pairs of QA were generated by machine translation of the original SQuAD1.0 and approximately 4000 pairs via crowdsourcing. In this study, we used two types of MRC models: rule-based baseline and advanced Transformer-based models. However, we have discovered that the latter outperforms the others; thus, we have decided to concentrate solely on Transformer-based architectures. Using XLMRoBERTa and multi-lingual BERT, we acquire an F1 score of 0.66 and 0.63, respectively.

📄 PDF Abstract BibTeX arXiv:2111.01543

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesMachine Reading ComprehensionMachine TranslationQuestion AnsweringTranslation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering

2024-05-21 · Hiba Maryam, Ling Fu, Jiajun Song, Tajrian ABM Shafayet 외

The development of Urdu scene text detection, recognition, and Visual Question Answering (VQA) technologies is crucial for advancing accessibility, information retrieval, and linguistic diversity in digital content, faci…

DiversityInformation RetrievalQuestion AnsweringRetrieval+4

UQA: Corpus for Urdu Question Answering

2024-05-02 · Samee Arif, Sualeha Farid, Awais Athar, Agha Ali Raza

This paper introduces UQA, a novel dataset for question answering and text comprehension in Urdu, a low-resource language with over 70 million native speakers. UQA is generated by translating the Stanford Question Answer…

Multilingual NLPQuestion AnsweringReading Comprehension

Multilingual Hematology Visual Question Answering Dataset

2026-06-24 · Hajra Malik, Hafiza Tooba Aftab, Abdul Rehman, Mohsen Ali 외 arxiv

Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for tasks such as Visual Question Answering. However, existing hematology …

Visual Question Answering

LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering

2024-10-16 · Faizan Faisal, Umair Yousaf

We present LEGAL-UQA, the first Urdu legal question-answering dataset derived from Pakistan's constitution. This parallel English-Urdu dataset includes 619 question-answer pairs, each with corresponding legal article con…

Optical Character Recognition (OCR)Question AnsweringRetrieval

UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking

2025-05-21 · Sarfraz Ahmad, Hasan Iqbal, Momina Ahsan, Numaan Naeem 외

The rapid use of large language models (LLMs) has raised critical concerns regarding the factual reliability of their outputs, especially in low-resource languages such as Urdu. Existing automated fact-checking solutions…

BenchmarkingClaim VerificationFact CheckingQuestion Answering+1