paper-with-me

홈 › Papers

AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data

2025-07-24 · Rana Alshaikh, Israa Alghanmi, Shelan Jeawak arxiv

The cognitive and reasoning abilities of large language models (LLMs) have enabled remarkable progress in natural language processing. However, their performance in interpreting structured data, especially in tabular formats, remains limited. Although benchmarks for English tabular data are widely available, Arabic is still underrepresented because of the limited availability of public resources and its unique language features. To address this gap, we present AraTable, a novel and comprehensive benchmark designed to evaluate the reasoning and understanding capabilities of LLMs when applied to Arabic tabular data. AraTable consists of various evaluation tasks, such as direct question answering, fact verification, and complex reasoning, involving a wide range of Arabic tabular sources. Our methodology follows a hybrid pipeline, where initial content is generated by LLMs and subsequently filtered and verified by human experts to ensure high dataset quality. Initial analyses using AraTable show that, while LLMs perform adequately on simpler tabular tasks such as direct question answering, they continue to face significant cognitive challenges when tasks require deeper reasoning and fact verification. This indicates that there are substantial opportunities for future work to improve performance on complex tabular reasoning tasks. We also propose a fully automated evaluation framework that uses a self-deliberation mechanism and achieves performance nearly identical to that of human judges. This research provides a valuable, publicly available resource and evaluation framework that can help accelerate the development of foundational models for processing and analysing Arabic structured data.

📄 PDF Abstract BibTeX arXiv:2507.18442

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringFact Verification

Similar Papers 제목 키워드 기반

Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues

2026-04-30 · Muhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan, Bilal Elbouardi 외 arxiv

There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standar…

Machine Translation

DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models

2025-10-31 · Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih 외 arxiv

We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluat…

AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects

2024-12-31 · Ahmad Mustapha, Hadi Al-Khansa, Hadi Al-Mubasher, Aya Mourad 외

Large Language Models (LLMs) have shown remarkable capabilities, not only in generating human-like text, but also in acquiring knowledge. This highlights the need to go beyond the typical Natural Language Processing down…

BenchmarkingMultiple-choice

Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks

2025-08-13 · Nouar AlDahoul, Yasir Zaki arxiv

Recent progress in large language models (LLMs) has showcased impressive proficiency in numerous Arabic natural language processing (NLP) applications. Nevertheless, their effectiveness in Arabic medical NLP domains has …

MizanQA: Benchmarking Large Language Models on Moroccan Legal Question Answering

2025-08-22 · Adil Bahaj, Mounir Ghogho arxiv

The rapid advancement of large language models (LLMs) has significantly propelled progress in natural language processing (NLP). However, their effectiveness in specialized, low-resource domains-such as Arabic legal cont…

Question AnsweringLegal Reasoning