paper-with-me

홈 › Papers

Urdu Morphology, Orthography and Lexicon Extraction

2022-04-06 · Muhammad Humayoun, Harald Hammarström, Aarne Ranta

Urdu is a challenging language because of, first, its Perso-Arabic script and second, its morphological system having inherent grammatical forms and vocabulary of Arabic, Persian and the native languages of South Asia. This paper describes an implementation of the Urdu language as a software API, and we deal with orthography, morphology and the extraction of the lexicon. The morphology is implemented in a toolkit called Functional Morphology (Forsberg & Ranta, 2004), which is based on the idea of dealing grammars as software libraries. Therefore this implementation could be reused in applications such as intelligent search of keywords, language training and infrastructure for syntax. We also present an implementation of a small part of Urdu syntax to demonstrate this reusability.

📄 PDF Abstract BibTeX arXiv:2204.03071

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Conventional Orthography for Dialectal Arabic

2012-05-01 · LREC 2012 5 · Nizar Habash, Mona Diab, Owen Rambow

Dialectal Arabic (DA) refers to the day-to-day vernaculars spoken in the Arab world. DA lives side-by-side with the official language, Modern Standard Arabic (MSA). DA differs from MSA on all levels of linguistic represe…

Speech Recognition

UniDic for Early Middle Japanese: a Dictionary for Morphological Analysis of Classical Japanese

2012-05-01 · LREC 2012 5 · Toshinobu Ogiso, Mamoru Komachi, Yasuharu Den, Yuji Matsumoto

In order to construct an annotated diachronic corpus of Japanese, we propose to create a new dictionary for morphological analysis of Early Middle Japanese (Classical Japanese) based on UniDic, a dictionary for Contempor…

Morphological Analysis

PronouncUR: An Urdu Pronunciation Lexicon Generator

2018-01-01 · LREC 2018 5 · Haris Bin Zia, Agha Ali Raza, Awais Athar

State-of-the-art speech recognition systems rely heavily on three basic components: an acoustic model, a pronunciation lexicon and a language model. To build these components, a researcher needs linguistic as well as tec…

Grapheme-to-Phoneme ConversionLanguage ModelingLanguage Modellingspeech-recognition+1

Computer-aided morphology expansion for Old Swedish

2014-05-01 · LREC 2014 5 · Yvonne Adesam, Malin Ahlberg, Peter Andersson, Gerlof Bouma 외

In this paper we describe and evaluate a tool for paradigm induction and lexicon extraction that has been applied to Old Swedish. The tool is semi-supervised and uses a small seed lexicon and unannotated corpora to deriv…

Learning Trilingual Dictionaries for Urdu -- Roman Urdu -- English

2019-08-01 · WS 2019 8 · Moiz Rauf, Sebastian Pad{\'o}

In this paper, we present an effort to generate a joint Urdu, Roman Urdu and English trilingual lexicon using automated methods. We make a case for using statistical machine translation approaches and parallel corpora fo…

Machine TranslationTranslationWord Alignment