paper-with-me

홈 › Papers

Tafsir Dataset: A Novel Multi-Task Benchmark for Named Entity Recognition and Topic Modeling in Classical Arabic Literature

2022-10-01 · COLING 2022 10 · Sajawel Ahmed, Rob van der Goot, Misbahur Rehman, Carl Kruse, Ömer Özsoy, Alexander Mehler, Gemma Roig

Various historical languages, which used to be lingua franca of science and arts, deserve the attention of current NLP research. In this work, we take the first data-driven steps towards this research line for Classical Arabic (CA) by addressing named entity recognition (NER) and topic modeling (TM) on the example of CA literature. We manually annotate the encyclopedic work of Tafsir Al-Tabari with span-based NEs, sentence-based topics, and span-based subtopics, thus creating the Tafsir Dataset with over 51,000 sentences, the first large-scale multi-task benchmark for CA. Next, we analyze our newly generated dataset, which we make open-source available, with current language models (lightweight BiLSTM, transformer-based MaChAmP) along a novel script compression method, thereby achieving state-of-the-art performance for our target task CA-NER. We also show that CA-TM from the perspective of historical topic models, which are central to Arabic studies, is very challenging. With this interdisciplinary work, we lay the foundations for future research on automatic analysis of CA literature.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERSentenceTopic Models

Similar Papers 제목 키워드 기반

Quranic Conversations: Developing a Semantic Search tool for the Quran using Arabic NLP Techniques

2023-11-09 · Yasser Shohoud, Maged Shoman, Sarah Abdelazim

The Holy Book of Quran is believed to be the literal word of God (Allah) as revealed to the Prophet Muhammad (PBUH) over a period of approximately 23 years. It is the book where God provides guidance on how to live a rig…

Business EthicsEthics

Towards A Time Based Video Search Engine for Al Quran Interpretation

2017-01-25 · Eljazzar Maged M., Hassan Afnan, AlSharkawy Amira A.

The number of Internet Muslim-users is remarkably increasing from all over the world countries. There are a lot of structured, and well-documented text resources for the Quran interpretation, Tafsir, over the Internet wi…

A Benchmark Dataset with Larger Context for Non-Factoid Question Answering over Islamic Text

2024-09-15 · Faiza Qamar, Seemab Latif, Rabia Latif

Accessing and comprehending religious texts, particularly the Quran (the sacred scripture of Islam) and Ahadith (the corpus of the sayings or traditions of the Prophet Muhammad), in today's digital era necessitates effic…

Question Answering

Transformer Tafsir at QIAS 2025 Shared Task: Hybrid Retrieval-Augmented Generation for Islamic Knowledge Question Answering

2025-09-28 · Muhammad Abu Ahmad, Mohamad Ballout, Raia Abu Ahmad, Elia Bruni arxiv

This paper presents our submission to the QIAS 2025 shared task on Islamic knowledge understanding and reasoning. We developed a hybrid retrieval-augmented generation (RAG) system that combines sparse and dense retrieval…

Question Answering

Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark

2026-04-14 · Terra Blevins, Stephen Mayhew, Marek Šuppa, Hila Gonen 외 arxiv

While multilingual language models promise to bring the benefits of LLMs to speakers of many languages, gold-standard evaluation benchmarks in most languages to interrogate these assumptions remain scarce. The Universal …