paper-with-me

Papers

A Systematic Benchmark of Machine Transliteration Models for the Tajik-Farsi Language Pair: A Comparative Study from Rule-Based to Transformer Architectures

2026-05-04 · Mullosharaf K. Arabov arxiv

This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic script) and Persian (Arabic script). A key contribution is the creation and validation of a unique parallel corpus aggregated from multiple heterogeneous sources, including crowdsourced projects, lexicographic pairs, parallel texts of "Shahnameh", diplomatic articles, texts of "Masnavi-i Ma'navi", official terminology lists, and transliterated correspondences. The initial dataset comprised 328,253 sentence pairs; a representative subset of 40,000 pairs was formed using stratified random sampling. The experiment compared six classes of models: rule-based baseline, LSTM with attention, character-level Transformer, G2P Transformer (trained from scratch), pre-trained multilingual models (mBART, mT5 with LoRA), and byte-level ByT5. Results demonstrate the overwhelming superiority of ByT5 (chrF++ 87.4 for Tajik to Farsi, 80.1 for reverse). The G2P Transformer significantly outperformed mBART (72.3 vs. 62.2 chrF++) despite limited data. Models using subword tokenization (mT5) failed completely (chrF++ less than 18.5). The findings demonstrate that for accurate transliteration of the Tajik-Farsi pair, architectures operating at the byte or character level are unequivocally more effective than traditional multilingual Seq2Seq models relying on subword tokenization.

📄 PDF Abstract BibTeX arXiv:2605.02270

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Connecting the Persian-speaking World through Transliteration

2025-02-27 · Rayyan Merchant, Akhilesh Kakolu Ramarao, Kevin Tang

Despite speaking mutually intelligible varieties of the same language, speakers of Tajik Persian, written in a modified Cyrillic alphabet, cannot read Iranian and Afghan texts written in the Perso-Arabic script. As the v…

Machine TranslationTransliteration

ParsTranslit: Truly Versatile Tajik-Farsi Transliteration

2025-10-08 · Rayyan Merchant, Kevin Tang arxiv

As a digraphic language, the Persian language utilizes two written standards: Perso-Arabic in Afghanistan and Iran, and Tajik-Cyrillic in Tajikistan. Despite the significant similarity between the dialects of each countr…

Tajik-Farsi Persian Transliteration Using Statistical Machine Translation

2012-05-01 · LREC 2012 5 · Chris Irwin Davis

Tajik Persian is a dialect of Persian spoken primarily in Tajikistan and written with a modified Cyrillic alphabet. Iranian Persian, or Farsi, as it is natively called, is the lingua franca of Iran and is written with th…

Machine TranslationTranslationTransliteration

Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus

2026-05-09 · Mullosharaf K. Arabov arxiv

This study addresses automatic transliteration from Tajik (Cyrillic script) to Persian (Perso-Arabic script). We present a curated, lexicographically verified parallel corpus of 52,152 Tajik--Persian words and short phra…

Challenges in Persian Electronic Text Analysis

2014-04-18 · Behrang QasemiZadeh, Saeed Rahimi, Mehdi Safaee Ghalati

Farsi, also known as Persian, is the official language of Iran and Tajikistan and one of the two main languages spoken in Afghanistan. Farsi enjoys a unified Arabic script as its writing system. In this paper we briefly …