paper-with-me

Papers

Multi-stage Training of Bilingual Islamic LLM for Neural Passage Retrieval

2025-01-17 · Vera Pavlova

This study examines the use of Natural Language Processing (NLP) technology within the Islamic domain, focusing on developing an Islamic neural retrieval model. By leveraging the robust XLM-R model, the research employs a language reduction technique to create a lightweight bilingual large language model (LLM). Our approach for domain adaptation addresses the unique challenges faced in the Islamic domain, where substantial in-domain corpora exist only in Arabic while limited in other languages, including English. The work utilizes a multi-stage training process for retrieval models, incorporating large retrieval datasets, such as MS MARCO, and smaller, in-domain datasets to improve retrieval performance. Additionally, we have curated an in-domain retrieval dataset in English by employing data augmentation techniques and involving a reliable Islamic source. This approach enhances the domain-specific dataset for retrieval, leading to further performance gains. The findings suggest that combining domain adaptation and a multi-stage training method for the bilingual Islamic neural retrieval model enables it to outperform monolingual models on downstream retrieval tasks.

📄 PDF Abstract BibTeX arXiv:2501.10175

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDomain AdaptationLanguage ModelingLanguage ModellingLarge Language ModelPassage RetrievalRetrievalXLM-R

Methods 이 논문이 사용한 방법론

XLM-R XLM-R

Similar Papers 제목 키워드 기반

From RAG to Agentic RAG for Faithful Islamic Question Answering

2026-01-12 · Gagan Bhatia, Hamdy Mubarak, Mustafa Jarrar, George Mikros 외 arxiv

Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences. Yet standard MCQ/MRC-style evaluations (MCQ: Multiple choice questio…

Machine Reading ComprehensionQuestion Answering

Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA

2026-03-09 · Ummar Abbas, Mourad Ouzzani, Mohamed Y. Eltabakh, Omar Sinan 외 arxiv

Large language models (LLMs) can answer religious knowledge queries fluently, yet they often hallucinate and misattribute sources, which is especially consequential in Islamic settings where users expect grounding in can…

QU-NLP at QIAS 2026: Multi-Stage QLoRA Fine-Tuning for Arabic Islamic Inheritance Reasoning

2026-03-29 · Mohammad AL-Smadi arxiv

Islamic inheritance law (ilm al-mawarıth) presents a challenging domain for evaluating large language models' structured reasoning capabilities, requiring multi-step legal analysis, rule-based blocking decisions, and pre…

Domain AdaptationLegal Reasoning

Constructing a Bilingual Hadith Corpus Using a Segmentation Tool

2020-05-01 · LREC 2020 5 · Shatha Altammami, Eric Atwell, Ammar Alsalka

This article describes the process of gathering and constructing a bilingual parallel corpus of Islamic Hadith, which is the set of narratives reporting different aspects of the prophet Muhammad{'}s life. The corpus data…

IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge

2026-03-24 · Ali Abdelaal, Mohammed Nader Al Haffar, Mahmoud Fawzi, Walid Magdy arxiv

Large language models are increasingly consulted for Islamic knowledge, yet no comprehensive benchmark evaluates their performance across core Islamic disciplines. We introduce IslamicMMLU, a benchmark of 10,013 multiple…

Bias Detection