paper-with-me

홈 › Papers

Sa‘7r: A Saudi Dialect Irony Dataset

2022-06-01 · OSACT (LREC) 2022 6 · Halah AlMazrua, Najla AlHazzani, Amaal AlDawod, Lama AlAwlaqi, Noura AlReshoudi, Hend Al-Khalifa, Luluh AlDhubayi

In sentiment analysis, detecting irony is considered a major challenge. The key problem with detecting irony is the difficulty to recognize the implicit and indirect phrases which signifies the opposite meaning. In this paper, we present Sa‘7r ساخرthe Saudi irony dataset, and describe our efforts in constructing it. The dataset was collected using Twitter API and it consists of 19,810 tweets, 8,089 of them are labeled as ironic tweets. We trained several models for irony detection task using machine learning models and deep learning models. The machine learning models include: K-Nearest Neighbor (KNN), Logistic Regression (LR), Support Vector Machine (SVM), and Naïve Bayes (NB). While the deep learning models include BiLSTM and AraBERT. The detection results show that among the tested machine learning models, the SVM outperformed other classifiers with an accuracy of 0.68. On the other hand, the deep learning models achieved an accuracy of 0.66 in the BiLSTM model and 0.71 in the AraBERT model. Thus, the AraBERT model achieved the most accurate result in detecting irony phrases in Saudi Dialect.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningSentiment Analysis

Similar Papers 제목 키워드 기반

SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora

2024-05-10 · Faisal Qarah

In this paper, we introduce SaudiBERT, a monodialect Arabic language model pretrained exclusively on Saudi dialectal text. To demonstrate the model's effectiveness, we compared SaudiBERT with six different multidialect A…

Language ModelingLanguage ModellingLarge Language ModelSentiment Analysis+2

Saudi-Dialect-ALLaM: LoRA Fine-Tuning for Dialectal Arabic Generation

2025-08-19 · Hassan Barmandah arxiv

Large language models (LLMs) for Arabic are still dominated by Modern Standard Arabic (MSA), with limited support for Saudi dialects such as Najdi and Hijazi. This underrepresentation hinders their ability to capture aut…

From Words to Proverbs: Evaluating LLMs Linguistic and Cultural Competence in Saudi Dialects with Absher

2025-07-14 · Renad Al-Monef, Hassan Alhuzali, Nora Alturayeif, Ashwag Alasmari arxiv

As large language models (LLMs) become increasingly central to Arabic NLP applications, evaluating their understanding of regional dialects and cultural nuances is essential, particularly in linguistically diverse settin…

DAICT: A Dialectal Arabic Irony Corpus Extracted from Twitter

2020-05-01 · LREC 2020 5 · Ines Abbes, Wajdi Zaghouani, Omaima El-Hardlo, Faten Ashour

Identifying irony in user-generated social media content has a wide range of applications; however to date Arabic content has received limited attention. To bridge this gap, this study builds a new open domain Arabic cor…

ArabicDialectHub: A Cross-Dialectal Arabic Learning Resource and Platform

2026-01-30 · Salem Lahlou arxiv

We present ArabicDialectHub, a cross-dialectal Arabic learning resource comprising 552 phrases across six varieties (Moroccan Darija, Lebanese, Syrian, Emirati, Saudi, and MSA) and an interactive web platform. Phrases we…

Distractor Generation