Extracting Morphophonology from Small Corpora
Probabilistic approaches have proven themselves well in learning phonological structure. In contrast, theoretical linguistics usually works with deterministic generalizations. The goal of this paper is to explore possible interactions between information-theoretic methods and deterministic linguistic knowledge and to examine some ways in which both can be used in tandem to extract phonological and morphophonological patterns from a small annotated dataset. Local and nonlocal processes in Mishar Tatar (Turkic/Kipchak) are examined as a case study.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Representation of Yine [Arawak] Morphology by Finite State Transducer Formalism
We represent the complexity of Yine (Arawak) morphology with a finite state transducer (FST) based morphological analyzer. Yine is a low-resource indigenous polysynthetic Peruvian language spoken by approximately 3,000 p…
Interactive Word Completion for Morphologically Complex Languages
Text input technologies for low-resource languages support literacy, content authoring, and language learning. However, tasks such as word completion pose a challenge for morphologically complex languages thanks to the c…
MORPHExtracting Mathematical Concepts from Text
We investigate different systems for extracting mathematical entities from English texts in the mathematical field of category theory as a first step for constructing a mathematical knowledge graph. We consider four diff…
Extracting RDF Triples from Raw Text
This manuscript manifests the results of our work on extracting RDF triples from raw text data. We took a corpus of news articles and applied several methods for extracting “subject - verb - object” relationships from te…
ArticlesFinite-state Model of Shupamem Reduplication
Shupamem, a language of Western Cameroon, is a tonal language which also exhibits the morpho-phonological process of full reduplication. This creates two challenges for finite-state model of its morpho-syntax and morphop…
model