paper-with-me

Papers

ORCA: A Challenging Benchmark for Arabic Language Understanding

2022-12-21 · AbdelRahim Elmadany, El Moatez Billah Nagoudi, Muhammad Abdul-Mageed

Due to their crucial role in all NLP, several benchmarks have been proposed to evaluate pretrained language models. In spite of these efforts, no public benchmark of diverse nature currently exists for evaluation of Arabic. This makes it challenging to measure progress for both Arabic and multilingual language models. This challenge is compounded by the fact that any benchmark targeting Arabic needs to take into account the fact that Arabic is not a single language but rather a collection of languages and varieties. In this work, we introduce ORCA, a publicly available benchmark for Arabic language understanding evaluation. ORCA is carefully constructed to cover diverse Arabic varieties and a wide range of challenging Arabic understanding tasks exploiting 60 different datasets across seven NLU task clusters. To measure current progress in Arabic NLU, we use ORCA to offer a comprehensive comparison between 18 multilingual and Arabic language models. We also provide a public leaderboard with a unified single-number evaluation metric (ORCA score) to facilitate future research.

📄 PDF Abstract BibTeX arXiv:2212.10758

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the importance of Data Scale in Pretraining Arabic Language Models

2024-01-15 · Abbas Ghaddar, Philippe Langlais, Mehdi Rezagholizadeh, Boxing Chen

Pretraining monolingual language models have been proven to be vital for performance in Arabic Natural Language Processing (NLP) tasks. In this paper, we conduct a comprehensive study on the role of data in Arabic Pretra…

DecoderLanguage ModelingLanguage Modelling

ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering

2026-03-02 · Aymen Lassoued, Mohamed Ali Souibgui, Yousri Kessentini arxiv

Document Visual Question Answering (DocVQA) remains challenging for existing Vision-Language Models (VLMs), especially under complex reasoning and multi-step workflows. Current approaches struggle to decompose intricate …

Visual Question Answering

D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning

2026-02-08 · Changli Tang, Tianyi Wang, Fengyun Rao, Jing Lyu 외 arxiv

Spoken dialogue is a primary source of information in videos; therefore, accurately identifying who spoke what and when is essential for deep video understanding. We introduce D-ORCA, a \textbf{d}ialogue-centric \textbf{…

Reinforcement LearningSpeaker IdentificationSpeech Recognition

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

2025-11-28 · Šimon Sedláček, Sara Barahona, Bolaji Yusuf, Laura Herrera-Alarcón 외 arxiv

Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reasoning and subjective tasks, they increas…

Question Answering

ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

2024-02-20 · Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman 외

The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Ar…

ArabicMMLULanguage Model EvaluationLanguage ModelingLanguage Modelling+2