paper-with-me

홈 › Papers

Building Arabic NLP from the Ground Up: Twenty Years of Lessons, Failures, and Open Problems

2026-05-20 · Wajdi Zaghouani arxiv

This paper reflects on twenty years of building NLP resources and research infrastructure for Arabic, a language spoken by hundreds of millions yet historically underserved relative to languages such as English or Chinese. The first decade focused on foundational linguistic infrastructure; the second shifted toward computational social science, social media analysis, and socially oriented applications. Rather than cataloguing outputs, the paper examines what the experience of building them revealed. Three counterintuitive lessons emerge: building datasets is as much a social process as a technical one; communities formed around shared tasks often matter more than the tasks themselves; and moving from language resources to computational social science exposes challenges that traditional NLP training does not address. We discuss three failures: a depression detection corpus that never reached clinical practice, a period of spreading across too many shared tasks without sufficient depth, and a long-standing assumption that Modern Standard Arabic infrastructure would transfer cleanly to dialectal tasks. These experiences suggest that the hardest problems in developing NLP for underserved communities are not linguistic but social, institutional, and epistemic, and require competencies the field rarely teaches.

📄 PDF Abstract BibTeX arXiv:2605.20786

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Survey on Arabic Named Entity Recognition: Past, Recent Advances, and Future Trends

2023-02-07 · Xiaoye Qu, Yingjie Gu, Qingrong Xia, Zechang Li 외

As more and more Arabic texts emerged on the Internet, extracting important information from these Arabic texts is especially useful. As a fundamental technology, Named entity recognition (NER) serves as the core compone…

Feature EngineeringLanguage ModelingLanguage Modellingnamed-entity-recognition+4

LCA and energy efficiency in buildings: mapping more than twenty years of research

2024-08-23 · F. Asdrubali, A. Fronzetti Colladon, L. Segneri, D. M. Gandola

Research on Life Cycle Assessment (LCA) is being conducted in various sectors, from analyzing building materials and components to comprehensive evaluations of entire structures. However, reviews of the existing literatu…

Faheem at NADI shared task: Identifying the dialect of Arabic tweet

2020-12-01 · COLING (WANLP) 2020 12 · Nouf AlShenaifi, Aqil Azmi

This paper describes Faheem (adj. of understand), our submission to NADI (Nuanced Arabic Dialect Identification) shared task. With so many Arabic dialects being under-studied due to the scarcity of the resources, the obj…

Dialect Identificationregression

Calliar: An Online Handwritten Dataset for Arabic Calligraphy

2021-06-20 · Zaid Alyafeai, Maged S. Al-shaibani, Mustafa Ghaleb, Yousif Ahmed Al-Wajih

Calligraphy is an essential part of the Arabic heritage and culture. It has been used in the past for the decoration of houses and mosques. Usually, such calligraphy is designed manually by experts with aesthetic insight…

Cultural Vocal Bursts Intensity PredictionDiversitySentence

Tarab: A Multi-Dialect Corpus of Arabic Lyrics and Poetry

2026-03-17 · Mo El-Haj arxiv

We introduce the Tarab Corpus, a large-scale cultural and linguistic resource that brings together Arabic song lyrics and poetry within a unified analytical framework. The corpus comprises 2.56 million verses and more th…