paper-with-me

Papers

Guidelines and Framework for a Large Scale Arabic Diacritized Corpus

2016-05-01 · LREC 2016 5 · Wajdi Zaghouani, Houda Bouamor, Abdelati Hawwari, Mona Diab, Ossama Obeid, Mahmoud Ghoneim, Sawsan Alqahtani, Kemal Oflazer

This paper presents the annotation guidelines developed as part of an effort to create a large scale manually diacritized corpus for various Arabic text genres. The target size of the annotated corpus is 2 million words. We summarize the guidelines and describe issues encountered during the training of the annotators. We also discuss the challenges posed by the complexity of the Arabic language and how they are addressed. Finally, we present the diacritization annotation procedure and detail the quality of the resulting annotations.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Proper Name Diacritization for Arabic Wikipedia: A Benchmark Dataset

2025-05-05 · Rawan Bondok, Mayar Nassar, Salam Khalifa, Kurt Micallef 외

Proper names in Arabic Wikipedia are frequently undiacritized, creating ambiguity in pronunciation and interpretation, especially for transliterated named entities of foreign origin. While transliteration and diacritizat…

Transliteration

YallaMorph: A Benchmark for Evaluating Arabic Morphological Generation in Large Language Models

2026-09-09 · Mahmoud Reda, Salam Khalifa, Reham Marzouk, Nizar Habash arxiv

Arabic morphology remains challenging for large language models, since fluent generation does not guarantee accurate morphosyntactic control. Existing Arabic evaluations mainly target downstream tasks and do not directly…

Constrained CTC Decoding for Efficient Diacritic Restoration

2026-07-21 · Rufael Marew, Amr Keleg, Hanan Aldarmaki arxiv

In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling fine-grained phonological distinctions. The speech modality has recently …

Maknuune: A Large Open Palestinian Arabic Lexicon

2022-10-24 · Shahd Dibas, Christian Khairallah, Nizar Habash, Omar Fayez Sadi 외

We present Maknuune, a large open lexicon for the Palestinian Arabic dialect. Maknuune has over 36K entries from 17K lemmas, and 3.7K roots. All entries include diacritized Arabic orthography, phonological transcription …

Take the Hint: Improving Arabic Diacritization with Partially-Diacritized Text

2023-06-06 · Parnia Bahar, Mattia Di Gangi, Nick Rossenbach, Mohammad Zeineldeen

Automatic Arabic diacritization is useful in many applications, ranging from reading support for language learners to accurate pronunciation predictor for downstream tasks like speech synthesis. While most of the previou…

Speech Synthesis