paper-with-me

Papers

ALUE: Arabic Language Understanding Evaluation

2021-04-01 · EACL (WANLP) 2021 4 · Haitham Seelawi, Ibraheem Tuffaha, Mahmoud Gzawi, Wael Farhan, Bashar Talafha, Riham Badawi, Zyad Sober, Oday Al-Dweik, Abed Alhakim Freihat, Hussein Al-Natsheh

The emergence of Multi-task learning (MTL)models in recent years has helped push thestate of the art in Natural Language Un-derstanding (NLU). We strongly believe thatmany NLU problems in Arabic are especiallypoised to reap the benefits of such models. Tothis end we propose the Arabic Language Un-derstanding Evaluation Benchmark (ALUE),based on 8 carefully selected and previouslypublished tasks. For five of these, we providenew privately held evaluation datasets to en-sure the fairness and validity of our benchmark.We also provide a diagnostic dataset to helpresearchers probe the inner workings of theirmodels.Our initial experiments show thatMTL models outperform their singly trainedcounterparts on most tasks. But in order to en-tice participation from the wider community,we stick to publishing singly trained baselinesonly. Nonetheless, our analysis reveals thatthere is plenty of room for improvement inArabic NLU. We hope that ALUE will playa part in helping our community realize someof these improvements. Interested researchersare invited to submit their results to our online,and publicly accessible leaderboard.

📄 PDF Abstract BibTeX

Code (1)

Alue-Benchmark/alue_baselines 공식 구현 pytorch

Tasks

DiagnosticFairnessMulti-Task Learning

Similar Papers 제목 키워드 기반

JABER and SABER: Junior and Senior Arabic BERt

2021-12-08 · Abbas Ghaddar, Yimeng Wu, Ahmad Rashid, Khalil Bibi 외

Language-specific pre-trained models have proven to be more accurate than multilingual ones in a monolingual evaluation setting, Arabic is no exception. However, we found that previously released Arabic BERT models were …

Language ModelingLanguage ModellingNER

JEEM: Vision-Language Understanding in Four Arabic Dialects

2025-03-27 · Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali, Abdelrahman Mohamed 외

We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. JEEM includes the tasks of image …

Image CaptioningQuestion AnsweringVisual Question Answering

Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Understanding

2022-05-21 · Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid 외

There is a growing body of work in recent years to develop pre-trained language models (PLMs) for the Arabic language. This work concerns addressing two major problems in existing Arabic PLMs which constraint progress of…

Natural Language Understanding

ORCA: A Challenging Benchmark for Arabic Language Understanding

2022-12-21 · AbdelRahim Elmadany, El Moatez Billah Nagoudi, Muhammad Abdul-Mageed

Due to their crucial role in all NLP, several benchmarks have been proposed to evaluate pretrained language models. In spite of these efforts, no public benchmark of diverse nature currently exists for evaluation of Arab…

CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks

2024-09-19 · Zhaozhi Qian, Faroq Altam, Muhammad Alqurishi, Riad Souissi

Large Language Models (LLMs) are the cornerstones of modern artificial intelligence systems. This paper introduces Juhaina, a Arabic-English bilingual LLM specifically designed to align with the values and preferences of…

Instruction FollowingOpen-Ended Question AnsweringQuestion Answering