paper-with-me

홈 › Papers

Cross-Lingual Learning within Arabic Script for Low-Resource HTR

2026-05-03 · Sana Al-azzawi, Elisa Barney, Marcus Liwicki arxiv

Handwritten Text Recognition (HTR) with limited labeled data remains a challenging problem, particularly for Arabic-script languages. Although modern sequence-based recognizers perform well in high-resource settings, their accuracy degrades sharply as training data becomes scarce. Arabic-script languages share a common writing system with substantial character overlap, motivating cross-lingual learning as a strategy to mitigate data scarcity. We conduct a controlled line-level study of cross-lingual joint training for Arabic-script HTR under low-resource regimes (number of samples K = 100, 500, 1000 labeled lines) on Arabic (KHATT), Urdu (NUST-UHWR) and Persian (PHTD). CRNN and Vision Transformer-based HTR-VT models are trained on the union of multiple related Arabic-script datasets to mitigate the data scarcity and are evaluated on individual target languages. Both architectures benefit from cross-language training under low-resource conditions. CRNN remains more effective under extremely limited target-language data, whereas the benefits of cross-language training for HTR-VT become less consistent as larger amounts of target-language data become available. On Persian (PHTD), joint training achieves a Character Error Rate (CER) of 9.99 , surpassing previously reported results despite not using the full available training data. On an additional Urdu dataset (UNHD), joint training reduces CER from 17.20 to 14.45.

📄 PDF Abstract BibTeX arXiv:2605.02089

Code (0)

등록된 구현이 없습니다.

Tasks

Handwritten Text Recognition

Similar Papers 제목 키워드 기반

Can Multilingual Language Models Transfer to an Unseen Dialect? A Case Study on North African Arabizi

2020-05-01 · Benjamin Muller, Benoit Sagot, Djamé Seddah

Building natural language processing systems for non standardized and low resource languages is a difficult challenge. The recent success of large-scale multilingual pretrained language models provides new modeling tools…

Dependency ParsingPart-Of-Speech TaggingTransliteration

Cross-Lingual Transfer from Related Languages: Treating Low-Resource Maltese as Multilingual Code-Switching

2024-01-30 · Kurt Micallef, Nizar Habash, Claudia Borg, Fadhl Eryani 외

Although multilingual language models exhibit impressive cross-lingual transfer capabilities on unseen languages, the performance on downstream tasks is impacted when there is a script disparity with the languages used i…

Cross-Lingual TransferTransliteration

EmoAra: Emotion-Preserving English Speech Transcription and Cross-Lingual Translation with Arabic Text-to-Speech

2026-02-01 · Besher Hassan, Ibrahim Alsarraj, Musaab Hasan, Yousef Melhim 외 arxiv

This work presents EmoAra, an end-to-end emotion-preserving pipeline for cross-lingual spoken communication, motivated by banking customer service where emotional context affects service quality. EmoAra integrates Speech…

Speech Emotion RecognitionEmotion ClassificationMachine TranslationSpeech Recognition

Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data

2025-09-16 · Kurt Micallef, Nizar Habash, Claudia Borg arxiv

Maltese is a unique Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. Despite its Semitic roots, its orthography is based on the Latin scri…

Machine TranslationData Augmentation

Training a Bilingual Language Model by Mapping Tokens onto a Shared Character Space

2024-02-25 · Aviad Rom, Kfir Bar

We train a bilingual Arabic-Hebrew language model using a transliterated version of Arabic texts in Hebrew, to ensure both languages are represented in the same script. Given the morphological, structural similarities, a…

Language ModelingLanguage ModellingMachine TranslationTranslation+1