paper-with-me

Papers

Open-Source Boundary-Annotated Corpus for Arabic Speech and Language Processing

2012-05-01 · LREC 2012 5 · Claire Brierley, Majdi Sawalha, Eric Atwell

A boundary-annotated and part-of-speech tagged corpus is a prerequisite for developing phrase break classifiers. Boundary annotations in English speech corpora are descriptive, delimiting intonation units perceived by the listener. We take a novel approach to phrase break prediction for Arabic, deriving our prosodic annotation scheme from Tajw{\=\i}d (recitation) mark-up in the Qur'an which we then interpret as additional text-based data for computational analysis. This mark-up is prescriptive, and signifies a widely-used recitation style, and one of seven original styles of transmission. Here we report on version 1.0 of our Boundary-Annotated Qur'an dataset of 77430 words and 8230 sentences, where each word is tagged with prosodic and syntactic information at two coarse-grained levels. In (Sawalha et al., 2012), we use the dataset in phrase break prediction experiments. This research is part of a larger-scale project to produce annotation schemes, language resources, algorithms, and applications for Classical and Modern Standard Arabic.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ChunkingDescriptiveSpeech SynthesisText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

OSIAN: Open Source International Arabic News Corpus - Preparation and Integration into the CLARIN-infrastructure

2019-08-01 · WS 2019 8 · Imad Zeroual, Dirk Goldhahn, Thomas Eckart, Abdelhak Lakhouaja

The World Wide Web has become a fundamental resource for building large text corpora. Broadcasting platforms such as news websites are rich sources of data regarding diverse topics and form a valuable foundation for rese…

ArticlesDescriptiveLEMMA

DAICT: A Dialectal Arabic Irony Corpus Extracted from Twitter

2020-05-01 · LREC 2020 5 · Ines Abbes, Wajdi Zaghouani, Omaima El-Hardlo, Faten Ashour

Identifying irony in user-generated social media content has a wide range of applications; however to date Arabic content has received limited attention. To bridge this gap, this study builds a new open domain Arabic cor…

Predicting Phrase Breaks in Classical and Modern Standard Arabic Text

2012-05-01 · LREC 2012 5 · Majdi Sawalha, Claire Brierley, Eric Atwell

We train and test two probabilistic taggers for Arabic phrase break prediction on a purpose-built, “gold standard”, boundary-annotated and PoS-tagged Qur'an corpus of 77430 words and 8230 sentences. In a related LREC pap…

ChunkingHuman ParsingPart-Of-Speech TaggingPOS+1

SALMA: Arabic Sense-Annotated Corpus and WSD Benchmarks

2023-10-29 · Mustafa Jarrar, Sanad Malaysha, Tymaa Hammouda, Mohammed Khalilia

SALMA, the first Arabic sense-annotated corpus, consists of ~34K tokens, which are all sense-annotated. The corpus is annotated using two different sense inventories simultaneously (Modern and Ghani). SALMA novelty lies …

Word Sense Disambiguation

A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment

2025-02-19 · Khalid N. Elmadani, Nizar Habash, Hanada Taha-Thomure

This paper introduces the Balanced Arabic Readability Evaluation Corpus BAREC, a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 68,182 sentences spanning 1+ million words, carefull…

Diversity