paper-with-me

Papers

Proper Name Diacritization for Arabic Wikipedia: A Benchmark Dataset

2025-05-05 · Rawan Bondok, Mayar Nassar, Salam Khalifa, Kurt Micallef, Nizar Habash

Proper names in Arabic Wikipedia are frequently undiacritized, creating ambiguity in pronunciation and interpretation, especially for transliterated named entities of foreign origin. While transliteration and diacritization have been well-studied separately in Arabic NLP,their intersection remains underexplored. In this paper, we introduce a new manually diacritized dataset of Arabic proper names of various origins with their English Wikipedia equivalent glosses, and present the challenges and guidelines we followed to create it. We benchmark GPT-4o on the task of recovering full diacritization given the undiacritized Arabic and English forms, and analyze its performance. Achieving 73% accuracy, our results underscore both the difficulty of the task and the need for improved models and resources. We release our dataset to facilitate further research on Arabic Wikipedia proper name diacritization.

📄 PDF Abstract BibTeX arXiv:2505.02656

Code (0)

등록된 구현이 없습니다.

Tasks

Transliteration

Similar Papers 제목 키워드 기반

Diacritization of Maghrebi Arabic Sub-Dialects

2018-10-15 · Ahmed Abdelali, Mohammed Attia, Younes Samih, Kareem Darwish 외

Diacritization process attempt to restore the short vowels in Arabic written text; which typically are omitted. This process is essential for applications such as Text-to-Speech (TTS). While diacritization of Modern Stan…

text-to-speechText to Speech

More Data, Fewer Diacritics: Scaling Arabic TTS

2026-03-02 · Ahmed Musleh, Yifan Zhang, Kareem Darwish arxiv

Arabic Text-to-Speech (TTS) research has been hindered by the availability of both publicly available training data and accurate Arabic diacritization models. In this paper, we address the limitation by exploring Arabic …

Speech RecognitionActivity Detection

Arabic Diacritization: Stats, Rules, and Hacks

2017-04-01 · WS 2017 4 · Kareem Darwish, Hamdy Mubarak, Ahmed Abdelali

In this paper, we present a new and fast state-of-the-art Arabic diacritizer that guesses the diacritics of words and then their case endings. We employ a Viterbi decoder at word-level with back-off to stem, morphologica…

DecoderPart-Of-Speech TaggingTransliterationWord Sense Disambiguation

Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need

2024-01-09 · Abderrahman Skiredj, Ismail Berrada

Automatic diacritization of Arabic text involves adding diacritical marks (diacritics) to the text. This task poses a significant challenge with noteworthy implications for computational processing and comprehension. In …

AllArabic Text Diacritizationtoken-classificationToken Classification+1

A System for Diacritizing Four Varieties of Arabic

2019-11-01 · IJCNLP 2019 11 · Hamdy Mubarak, Ahmed Abdelali, Kareem Darwish, Mohamed Eldesouki 외

Short vowels, aka diacritics, are more often omitted when writing different varieties of Arabic including Modern Standard Arabic (MSA), Classical Arabic (CA), and Dialectal Arabic (DA). However, diacritics are required t…

Feature Engineeringtext-to-speechText to Speech