A Rule-based Kurdish Text Transliteration System
In this article, we present a rule-based approach for transliterating two mostly used orthographies in Sorani Kurdish. Our work consists of detecting a character in a word by removing the possible ambiguities and mapping it into the target orthography. We describe different challenges in Kurdish text mining and propose novel ideas concerning the transliteration task for Sorani Kurdish. Our transliteration system, named Wergor, achieves 82.79% overall precision and more than 99% in detecting the double-usage characters. We also present a manually transliterated corpus for Kurdish.
Code (1)
Tasks
TransliterationSimilar Papers 제목 키워드 기반
Transliterating Kurdish texts in Latin into Persian-Arabic script
Kurdish is written in different scripts. The two most popular scripts are Latin and Persian-Arabic. However, not all Kurdish readers are familiar with both mentioned scripts that could be resolved by automatic transliter…
TransliterationLeveraging Multilingual News Websites for Building a Kurdish Parallel Corpus
Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation systems thanks to deep learning methods, p…
ArticlesMachine TranslationTranslationTransliterationKLPT – Kurdish Language Processing Toolkit
Despite the recent advances in applying language-independent approaches to various natural language processing tasks thanks to artificial intelligence, some language-specific tools are still essential to process a langua…
DiversityLemmatizationTransliterationBuilding a Lemmatizer and a Spell-checker for Sorani Kurdish
The present paper aims at presenting a lemmatization and a word-level error correction system for Sorani Kurdish. We propose a hybrid approach based on the morphological rules and a n-gram language model. We have called …
Language ModelingLanguage ModellingLemmatizationHunspell for Sorani Kurdish Spell Checking and Morphological Analysis
Spell checking and morphological analysis are two fundamental tasks in text and natural language processing and are addressed in the early stages of the development of language technology. Despite the previous efforts, t…
Morphological Analysis