paper-with-me

홈 › Papers

Transliteration and alignment of parallel texts from Cyrillic to Latin

2014-05-01 · LREC 2014 5 · Mircea Petic, Daniela G{\^\i}fu

This article describes a methodology of recovering and preservation of old Romanian texts and problems related to their recognition. Our focus is to create a gold corpus for Romanian language (the novella Sania), for both alphabets used in Transnistria ― Cyrillic and Latin. The resource is available for similar researches. This technology is based on transliteration and semiautomatic alignment of parallel texts at the level of letter/lexem/multiwords. We have analysed every text segment present in this corpus and discovered other conventions of writing at the level of transliteration, academic norms and editorial interventions. These conventions allowed us to elaborate and implement some new heuristics that make a correct automatic transliteration process. Sometimes the words of Latin script are modified in Cyrillic script from semantic reasons (for instance, editor{'}s interpretation). Semantic transliteration is seen as a good practice in introducing multiwords from Cyrillic to Latin. Not only does it preserve how a multiwords sound in the source script, but also enables the translator to modify in the original text (here, choosing the most common sense of an expression). Such a technology could be of interest to lexicographers, but also to specialists in computational linguistics to improve the actual transliteration standards.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningMachine TranslationTransliteration

Similar Papers 제목 키워드 기반

Uzbek Cyrillic-Latin-Cyrillic Machine Transliteration

2021-01-13 · B. Mansurov, A. Mansurov

In this paper, we introduce a data-driven approach to transliterating Uzbek dictionary words from the Cyrillic script into the Latin script, and vice versa. We heuristically align characters of words in the source script…

Transliteration

Transliteration for Low-Resource Code-Switching Texts: Building an Automatic Cyrillic-to-Latin Converter for Tatar

2021-06-01 · NAACL (CALCS) 2021 6 · Chihiro Taguchi, Yusuke Sakai, Taro Watanabe

We introduce a Cyrillic-to-Latin transliterator for the Tatar language based on subword-level language identification. The transliteration is a challenging task due to the following two reasons. First, because modern Tat…

Language IdentificationTransliteration

A machine transliteration tool between Uzbek alphabets

2022-05-19 · Ulugbek Salaev, Elmurod Kuriyozov, Carlos Gómez-Rodríguez

Machine transliteration, as defined in this paper, is a process of automatically transforming written script of words from a source alphabet into words of another target alphabet within the same language, while preservin…

Transliteration

Connecting the Persian-speaking World through Transliteration

2025-02-27 · Rayyan Merchant, Akhilesh Kakolu Ramarao, Kevin Tang

Despite speaking mutually intelligible varieties of the same language, speakers of Tajik Persian, written in a modified Cyrillic alphabet, cannot read Iranian and Afghan texts written in the Perso-Arabic script. As the v…

Machine TranslationTransliteration

A Systematic Benchmark of Machine Transliteration Models for the Tajik-Farsi Language Pair: A Comparative Study from Rule-Based to Transformer Architectures

2026-05-04 · Mullosharaf K. Arabov arxiv

This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic script) and Persian (Arabic script). A key contribution is the creatio…