paper-with-me

Papers

Graphemic ambiguous queries on Arabic-scripted historical corpora

2019-09-01 · RANLP 2019 9 · Alicia Gonz{\'a}lez Mart{\'\i}nez
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Graphemic Normalization of the Perso-Arabic Script

2022-10-21 · Raiomond Doctor, Alexander Gutkin, Cibu Johny, Brian Roark 외

Since its original appearance in 1991, the Perso-Arabic script representation in Unicode has grown from 169 to over 440 atomic isolated characters spread over several code pages representing standard letters, various dia…

Language ModelingLanguage ModellingMachine Translation

Automatic Long Audio Alignment and Confidence Scoring for Conversational Arabic Speech

2014-05-01 · LREC 2014 5 · Mohamed Elmahdy, Mark Hasegawa-Johnson, Eiman Mustafawi

In this paper, a framework for long audio alignment for conversational Arabic speech is proposed. Accurate alignments help in many speech processing tasks such as audio indexing, speech recognizer acoustic model (AM) tra…

Language Modellingspeech-recognitionSpeech Recognition

A Tale of Two Scripts: Transliteration and Post-Correction for Judeo-Arabic

2025-07-07 · Juan Moreno Gonzalez, Bashar Alhafni, Nizar Habash arxiv

Judeo-Arabic refers to Arabic variants historically spoken by Jewish communities across the Arab world, primarily during the Middle Ages. Unlike standard Arabic, it is written in Hebrew script by Jewish writers and for J…

Machine Translation

Investigating Diatopic Variation in a Historical Corpus

2017-04-01 · WS 2017 4 · Stefanie Dipper, S Waldenberger, ra

This paper investigates diatopic variation in a historical corpus of German. Based on equivalent word forms from different language areas, replacement rules and mappings are derived which describe the relations between t…

AraMS-28k: The Largest Publicly Released Line-Level Dataset of Historical Arabic Manuscripts with Margin and Insertion-Anchor Annotations

2026-08-27 · Mohamed Guechaoui, Mohamed Diaa Zellagui, Souleyman Chaib, Sahraoui Dhelim arxiv

We introduce AraMS-28k, the largest publicly released line-level dataset of genuine historical Arabic manuscripts, comprising 14 books, 3,043 pages, and 28,600 annotated text lines (27,971 main-text, 629 margin). Thirtee…