paper-with-me

Papers

Cross-Language Approach for Quranic QA

2025-01-29 · Islam Oshallah, Mohamed Basem, Ali Hamdi, Ammar Mohammed

Question answering systems face critical limitations in languages with limited resources and scarce data, making the development of robust models especially challenging. The Quranic QA system holds significant importance as it facilitates a deeper understanding of the Quran, a Holy text for over a billion people worldwide. However, these systems face unique challenges, including the linguistic disparity between questions written in Modern Standard Arabic and answers found in Quranic verses written in Classical Arabic, and the small size of existing datasets, which further restricts model performance. To address these challenges, we adopt a cross-language approach by (1) Dataset Augmentation: expanding and enriching the dataset through machine translation to convert Arabic questions into English, paraphrasing questions to create linguistic diversity, and retrieving answers from an English translation of the Quran to align with multilingual training requirements; and (2) Language Model Fine-Tuning: utilizing pre-trained models such as BERT-Medium, RoBERTa-Base, DeBERTa-v3-Base, ELECTRA-Large, Flan-T5, Bloom, and Falcon to address the specific requirements of Quranic QA. Experimental results demonstrate that this cross-language approach significantly improves model performance, with RoBERTa-Base achieving the highest MAP@10 (0.34) and MRR (0.52), while DeBERTa-v3-Base excels in Recall@10 (0.50) and Precision@10 (0.24). These findings underscore the effectiveness of cross-language strategies in overcoming linguistic barriers and advancing Quranic QA systems

📄 PDF Abstract BibTeX arXiv:2501.17449

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationQuestion AnsweringTranslation

Methods 이 논문이 사용한 방법론

ADOPT Please enter a description about the method here
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Flan-T5 Flan-T5 is the instruction fine-tuned version of T5 or Text-to-Text Transfer Transformer Language Model.

Similar Papers 제목 키워드 기반

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

2026-06-18 · Nabil Mosharraf Hossain, Riasat Islam, Unaizah Obaidellah arxiv

Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines. However, existing ASR models often exhibit high Wo…

Self-Supervised LearningSpeech Recognition

A computational system to handle the orthographic layer of tajwid in contemporary Quranic Orthography

2025-05-16 · Alicia González Martínez

Contemporary Quranic Orthography (CQO) relies on a precise system of phonetic notation that can be traced back to the early stages of Islam, when the Quran was mainly oral in nature and the first written renderings of it…

Quran-MD: A Fine-Grained Multilingual Multimodal Dataset of the Quran

2026-01-25 · Muhammad Umar Salman, Mohammad Areeb Qazi, Mohammed Talha Alam arxiv

We present Quran MD, a comprehensive multimodal dataset of the Quran that integrates textual, linguistic, and audio dimensions at the verse and word levels. For each verse (ayah), the dataset provides its original Arabic…

Text-To-Speech SynthesisSemantic RetrievalSpeech RecognitionStyle Transfer

Quranic Audio Dataset: Crowdsourced and Labeled Recitation from Non-Arabic Speakers

2024-05-04 · Raghad Salameh, Mohamad Al Mdfaa, Nursultan Askarbekuly, Manuel Mazzara

This paper addresses the challenge of learning to recite the Quran for non-Arabic speakers. We explore the possibility of crowdsourcing a carefully annotated Quranic dataset, on top of which AI models can be built to sim…

Tadabur: A Large-Scale Quran Audio Dataset

2026-04-21 · Faisal Alherran arxiv

Despite growing interest in Quranic data research, existing Quran datasets remain limited in both scale and diversity. To address this gap, we present Tadabur, a large-scale Quran audio dataset. Tadabur comprises more th…