paper-with-me

홈 › Papers

JobArabi: An Arabic Corpus and Analysis of Job Announcements from Social Media

2026-05-20 · Wajdi Zaghouani, Shimaa Amer Ibrahim, Mabrouka Bessghaier, Houda Bouamor arxiv

This paper introduces JobArabi, a large-scale corpus of Arabic job announcements collected from social media between January 2024 and October 2025. The dataset contains 20,528 public posts from X and captures more than two years of employment-related discourse across Arabic-speaking online communities. The corpus was compiled using a linguistically informed query framework covering 21 Arabic keyword families that reflect gendered, plural, formal, and dialectal expressions of recruitment language. The resulting dataset includes posts from institutional, commercial, and individual accounts and provides metadata such as timestamps, engagement indicators, and geolocation when available, enabling temporal and regional analysis of employment discourse. Quantitative analysis reveals several sociolinguistic patterns in online recruitment, including the persistence of gendered hiring language, regional variation in occupational demand, and the emotional framing of recruitment messages. These findings highlight the potential of Arabic social media as a resource for studying labor market communication and linguistic change. The JobArabi corpus, together with documentation and collection scripts, will be released to support research in Arabic NLP, computational social science, and digital labor studies.

📄 PDF Abstract BibTeX arXiv:2605.20960

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Arabic Tweets Sentiment Analysis Dataset (ATSAD) using Distant Supervision and Self Training

2020-05-01 · LREC 2020 5 · Kathrein Abu Kwaik, Stergios Chatzikyriakidis, Simon Dobnik, Motaz Saad 외

As the number of social media users increases, they express their thoughts, needs, socialise and publish their opinions reviews. For good social media sentiment analysis, good quality resources are needed, and the lack o…

8kArabic Sentiment AnalysisSentiment Analysis

Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus

2026-05-21 · Wajdi Zaghouani, Mabrouka Bessghaier, MD. Rafiul Biswas, Shimaa Amer Ibrahim arxiv

This paper presents the Arabic Women and Society Corpus, a ten year collection of 252,487 public Arabic Facebook posts related to women's empowerment and social wellbeing. The corpus was collected from 51,660 pages acros…

Language Identification

Classifying Arabic dialect text in the Social Media Arabic Dialect Corpus (SMADC)

2019-07-01 · WS 2019 7 · Areej Alshutayri, Eric Atwell

DAICT: A Dialectal Arabic Irony Corpus Extracted from Twitter

2020-05-01 · LREC 2020 5 · Ines Abbes, Wajdi Zaghouani, Omaima El-Hardlo, Faten Ashour

Identifying irony in user-generated social media content has a wide range of applications; however to date Arabic content has received limited attention. To bridge this gap, this study builds a new open domain Arabic cor…

SentiALG: Automated Corpus Annotation for Algerian Sentiment Analysis

2018-08-15 · Imane Guellil, Ahsan Adeel, Faical Azouaou, Amir Hussain

Data annotation is an important but time-consuming and costly procedure. To sort a text into two classes, the very first thing we need is a good annotation guideline, establishing what is required to qualify for each cla…

Sentiment AnalysisTransliteration