Urdu Morphology, Orthography and Lexicon Extraction
Urdu is a challenging language because of, first, its Perso-Arabic script and second, its morphological system having inherent grammatical forms and vocabulary of Arabic, Persian and the native languages of South Asia. This paper describes an implementation of the Urdu language as a software API, and we deal with orthography, morphology and the extraction of the lexicon. The morphology is implemented in a toolkit called Functional Morphology (Forsberg & Ranta, 2004), which is based on the idea of dealing grammars as software libraries. Therefore this implementation could be reused in applications such as intelligent search of keywords, language training and infrastructure for syntax. We also present an implementation of a small part of Urdu syntax to demonstrate this reusability.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Conventional Orthography for Dialectal Arabic
Dialectal Arabic (DA) refers to the day-to-day vernaculars spoken in the Arab world. DA lives side-by-side with the official language, Modern Standard Arabic (MSA). DA differs from MSA on all levels of linguistic represe…
Speech RecognitionUniDic for Early Middle Japanese: a Dictionary for Morphological Analysis of Classical Japanese
In order to construct an annotated diachronic corpus of Japanese, we propose to create a new dictionary for morphological analysis of Early Middle Japanese (Classical Japanese) based on UniDic, a dictionary for Contempor…
Morphological AnalysisPronouncUR: An Urdu Pronunciation Lexicon Generator
State-of-the-art speech recognition systems rely heavily on three basic components: an acoustic model, a pronunciation lexicon and a language model. To build these components, a researcher needs linguistic as well as tec…
Grapheme-to-Phoneme ConversionLanguage ModelingLanguage Modellingspeech-recognition+1Computer-aided morphology expansion for Old Swedish
In this paper we describe and evaluate a tool for paradigm induction and lexicon extraction that has been applied to Old Swedish. The tool is semi-supervised and uses a small seed lexicon and unannotated corpora to deriv…
Learning Trilingual Dictionaries for Urdu -- Roman Urdu -- English
In this paper, we present an effort to generate a joint Urdu, Roman Urdu and English trilingual lexicon using automated methods. We make a case for using statistical machine translation approaches and parallel corpora fo…
Machine TranslationTranslationWord Alignment