Graphemic ambiguous queries on Arabic-scripted historical corpora
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Graphemic Normalization of the Perso-Arabic Script
Since its original appearance in 1991, the Perso-Arabic script representation in Unicode has grown from 169 to over 440 atomic isolated characters spread over several code pages representing standard letters, various dia…
Language ModelingLanguage ModellingMachine TranslationAutomatic Long Audio Alignment and Confidence Scoring for Conversational Arabic Speech
In this paper, a framework for long audio alignment for conversational Arabic speech is proposed. Accurate alignments help in many speech processing tasks such as audio indexing, speech recognizer acoustic model (AM) tra…
Language Modellingspeech-recognitionSpeech RecognitionA Tale of Two Scripts: Transliteration and Post-Correction for Judeo-Arabic
Judeo-Arabic refers to Arabic variants historically spoken by Jewish communities across the Arab world, primarily during the Middle Ages. Unlike standard Arabic, it is written in Hebrew script by Jewish writers and for J…
Machine TranslationInvestigating Diatopic Variation in a Historical Corpus
This paper investigates diatopic variation in a historical corpus of German. Based on equivalent word forms from different language areas, replacement rules and mappings are derived which describe the relations between t…
AraMS-28k: The Largest Publicly Released Line-Level Dataset of Historical Arabic Manuscripts with Margin and Insertion-Anchor Annotations
We introduce AraMS-28k, the largest publicly released line-level dataset of genuine historical Arabic manuscripts, comprising 14 books, 3,043 pages, and 28,600 annotated text lines (27,971 main-text, 629 margin). Thirtee…