paper-with-me

Papers

Transliteration for Low-Resource Code-Switching Texts: Building an Automatic Cyrillic-to-Latin Converter for Tatar

2021-06-01 · NAACL (CALCS) 2021 6 · Chihiro Taguchi, Yusuke Sakai, Taro Watanabe

We introduce a Cyrillic-to-Latin transliterator for the Tatar language based on subword-level language identification. The transliteration is a challenging task due to the following two reasons. First, because modern Tatar texts often contain intra-word code-switching to Russian, a different transliteration set of rules needs to be applied to each morpheme depending on the language, which necessitates morpheme-level language identification. Second, the fact that Tatar is a low-resource language, with most of the texts in Cyrillic, makes it difficult to prepare a sufficient dataset. Given this situation, we proposed a transliteration method based on subword-level language identification. We trained a language classifier with monolingual Tatar and Russian texts, and applied different transliteration rules in accord with the identified language. The results demonstrate that our proposed method outscores other Tatar transliteration tools, and imply that it correctly transcribes Russian loanwords to some extent.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationTransliteration

Similar Papers 제목 키워드 기반

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety

2026-07-11 · Chigozirim Ifebi, Brent Kong, Ayushi Mehrotra arxiv

Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a m…

Part-of-Speech Tagging for Code-Switched, Transliterated Texts without Explicit Language Identification

2018-10-01 · EMNLP 2018 10 · Kelsey Ball, Dan Garrette

Code-switching, the use of more than one language within a single utterance, is ubiquitous in much of the world, but remains a challenge for NLP largely due to the lack of representative data for training models. In this…

Language IdentificationPart-Of-Speech TaggingTransliterationWord Embeddings

Balanced End-to-End Monolingual pre-training for Low-Resourced Indic Languages Code-Switching Speech Recognition

2021-06-10 · Amir Hussein, Shammur Chowdhury, Najim Dehak, Ahmed Ali

The success in designing Code-Switching (CS) ASR often depends on the availability of the transcribed CS resources. Such dependency harms the development of ASR in low-resourced languages such as Bengali and Hindi. In th…

Language Modellingspeech-recognitionSpeech RecognitionTransfer Learning+1

Universal Dependency Parsing for Hindi-English Code-switching

2018-04-16 · NAACL 2018 6 · Irshad Ahmad Bhat, Riyaz Ahmad Bhat, Manish Shrivastava, Dipti Misra Sharma

Code-switching is a phenomenon of mixing grammatical structures of two or more languages under varied social constraints. The code-switching data differ so radically from the benchmark corpora used in NLP community that …

Dependency ParsingLanguage IdentificationTAGTransliteration

A Tale of Two Scripts: Transliteration and Post-Correction for Judeo-Arabic

2025-07-07 · Juan Moreno Gonzalez, Bashar Alhafni, Nizar Habash arxiv

Judeo-Arabic refers to Arabic variants historically spoken by Jewish communities across the Arab world, primarily during the Middle Ages. Unlike standard Arabic, it is written in Hebrew script by Jewish writers and for J…

Machine Translation