Taqyim: Evaluating Arabic NLP Tasks Using ChatGPT Models
Large language models (LLMs) have demonstrated impressive performance on various downstream tasks without requiring fine-tuning, including ChatGPT, a chat-based model built on top of LLMs such as GPT-3.5 and GPT-4. Despite having a lower training proportion compared to English, these models also exhibit remarkable capabilities in other languages. In this study, we assess the performance of GPT-3.5 and GPT-4 models on seven distinct Arabic NLP tasks: sentiment analysis, translation, transliteration, paraphrasing, part of speech tagging, summarization, and diacritization. Our findings reveal that GPT-4 outperforms GPT-3.5 on five out of the seven tasks. Furthermore, we conduct an extensive analysis of the sentiment analysis task, providing insights into how LLMs achieve exceptional results on a challenging dialectal dataset. Additionally, we introduce a new Python interface https://github.com/ARBML/Taqyim that facilitates the evaluation of these tasks effortlessly.
Code (1)
Tasks
Part-Of-Speech TaggingSentiment AnalysisTransliterationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP
ChatGPT's emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks. However, the model's efficacy across diverse linguistic contexts remains …
Natural Language UnderstandingCan ChatGPT capture swearing nuances? Evidence from translating Arabic oaths
This study sets out to answer one major question: Can ChatGPT capture swearing nuances? It presents an empirical study on the ability of ChatGPT to translate Arabic oath expressions into English. 30 Arabic oath expressio…
TranslationThe Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic
Despite the growing importance of Arabic as a global language, there is a notable lack of language models pre-trained exclusively on Arabic data. This shortage has led to limited benchmarks available for assessing langua…
Language ModelingLanguage ModellingMathematical ReasoningAraSpider: Democratizing Arabic-to-SQL
This study presents AraSpider, the first Arabic version of the Spider dataset, aimed at improving natural language processing (NLP) in the Arabic-speaking community. Four multilingual translation models were tested for t…
Text to SQLText-To-SQLTranslationTARJAMAT: Evaluation of Bard and ChatGPT on Machine Translation of Ten Arabic Varieties
Despite the purported multilingual proficiency of instruction-finetuned large language models (LLMs) such as ChatGPT and Bard, the linguistic inclusivity of these models remains insufficiently explored. Considering this …
Dialogue GenerationMachine TranslationQuestion AnsweringTranslation