paper-with-me

홈 › Papers

Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation

2026-04-03 · Leen AlQadi, Ahmed Alzubaidi, Mohammed Alyafeai, Hamza Alobeidli, Maitha Alhammadi, Shaikha Alsuwaidi, Omar Alkaabi, Basma El Amel Boussaha, Hakim Hacid arxiv

We present QIMMA, a quality-assured Arabic LLM leaderboard that places systematic benchmark validation at its core. Rather than aggregating existing resources as-is, QIMMA applies a multi-model assessment pipeline combining automated LLM judgment with human review to surface and resolve systematic quality issues in well-established Arabic benchmarks before evaluation. The result is a curated, multi-domain, multi-task evaluation suite of over 52k samples, grounded predominantly in native Arabic content; code evaluation tasks are the sole exception, as they are inherently language-agnostic. Transparent implementation via LightEval, EvalPlus and public release of per-sample inference outputs make QIMMA a reproducible and community-extensible foundation for Arabic NLP evaluation.

📄 PDF Abstract BibTeX arXiv:2604.03395

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

3LM: Bridging Arabic, STEM, and Code through Benchmarking

2025-07-21 · Basma El Amel Boussaha, Leen AlQadi, Mugariya Farooq, Shaikha Alsuwaidi 외 arxiv

Arabic is one of the most widely spoken languages in the world, yet efforts to develop and evaluate Large Language Models (LLMs) for Arabic remain relatively limited. Most existing Arabic benchmarks focus on linguistic, …

Code Generation

Islamic Large Language Models: From Knowledge Acquisition to Trustworthy and Hallucination-Resistant AI

2026-06-15 · Mohammed Amine Mouhoub arxiv

Large language models (LLMs) are increasingly used for knowledge-intensive question answering, including religious and legal questions. Islamic knowledge is a particularly demanding setting: answers are expected to be gr…

Question AnsweringLegal Reasoning

Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale

2025-09-17 · Hasan Abed Al Kader Hammoud, Mohammad Zbeeb, Bernard Ghanem arxiv

We present Hala, a family of Arabic-centric instruction and translation models built with our translate-and-tune pipeline. We first compress a strong AR$\leftrightarrow$EN teacher to FP8 (yielding $\sim$2$\times$ higher …

Instruction Following

Arabic Prompts with English Tools: A Benchmark

2026-01-08 · Konstantin Kubrak, Ahmed El-Moselhy, Ammar Alsulami, Remaz Altuwaim 외 arxiv

Large Language Models (LLMs) are now integral to numerous industries, increasingly serving as the core reasoning engine for autonomous agents that perform complex tasks through tool-use. While the development of Arabic-n…

MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks

2025-05-06 · Mouath Abu Daoud, Chaimae Abouzahir, Leen Kharouf, Walid Al-Eisawi 외

Large Language Models (LLMs) have demonstrated significant promise for various applications in healthcare. However, their efficacy in the Arabic medical domain remains unexplored due to the lack of high-quality domain-sp…

BenchmarkingMultiple-choiceQuestion Answering