AmQA: Amharic Question Answering Dataset
Question Answering (QA) returns concise answers or answer lists from natural language text given a context document. Many resources go into curating QA datasets to advance robust models' development. There is a surge of QA datasets for languages like English, however, this is not true for Amharic. Amharic, the official language of Ethiopia, is the second most spoken Semitic language in the world. There is no published or publicly available Amharic QA dataset. Hence, to foster the research in Amharic QA, we present the first Amharic QA (AmQA) dataset. We crowdsourced 2628 question-answer pairs over 378 Wikipedia articles. Additionally, we run an XLMR Large-based baseline model to spark open-domain QA research interest. The best-performing baseline achieves an F-score of 69.58 and 71.74 in reader-retriever QA and reading comprehension settings respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesQuestion AnsweringReading ComprehensionSimilar Papers 제목 키워드 기반
AmaSQuAD: A Benchmark for Amharic Extractive Question Answering
This research presents a novel framework for translating extractive question-answering datasets into low-resource languages, as demonstrated by the creation of the AmaSQuAD dataset, a translation of SQuAD 2.0 into Amhari…
Extractive Question-AnsweringQuestion AnsweringXLM-RAMQA: An Adversarial Dataset for Benchmarking Bias of LLMs in Medicine and Healthcare
Large language models (LLMs) are reaching expert-level accuracy on medical diagnosis questions, yet their mistakes and the biases behind them pose life-critical risks. Bias linked to race, sex, and socioeconomic status i…
BenchmarkingMedical DiagnosisMedical Question AnsweringQuestion AnsweringWhat are the limits of cross-lingual dense passage retrieval for low-resource languages?
In this paper, we analyze the capabilities of the multi-lingual Dense Passage Retriever (mDPR) for extremely low-resource languages. In the Cross-lingual Open-Retrieval Answer Generation (CORA) pipeline, mDPR achieves su…
Answer GenerationLanguage ModelingLanguage ModellingPassage Retrieval+2Question Answering Classification for Amharic Social Media Community Based Questions
In this work, we build a Question Answering (QA) classification dataset from a social media platform, namely the Telegram public channel called @AskAnythingEthiopia. The channel has more than 78k subscribers and has exis…
8kQuestion AnsweringTransliterationAdvancing Question Answering on Handwritten Documents: A State-of-the-Art Recognition-Based Model for HW-SQuAD
Question-answering handwritten documents is a challenging task with numerous real-world applications. This paper proposes a novel recognition-based approach that improves upon the previous state-of-the-art on the HW-SQuA…
Question AnsweringRetrieval