Comparing morphological complexity of Spanish, Otomi and Nahuatl
We use two small parallel corpora for comparing the morphological complexity of Spanish, Otomi and Nahuatl. These are languages that belong to different linguistic families, the latter are low-resourced. We take into account two quantitative criteria, on one hand the distribution of types over tokens in a corpus, on the other, perplexity and entropy as indicators of word structure predictability. We show that a language can be complex in terms of how many different morphological word forms can produce, however, it may be less complex in terms of predictability of its internal structure of words.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
BPE vs. Morphological Segmentation: A Case Study on Machine Translation of Four Polysynthetic Languages
Morphologically-rich polysynthetic languages present a challenge for NLP systems due to data sparsity, and a common strategy to handle this issue is to apply subword segmentation. We investigate a wide variety of supervi…
Machine TranslationSegmentationTranslationCPLM, a Parallel Corpus for Mexican Languages: Development and Interface
Mexico is a Spanish speaking country that has a great language diversity, with 68 linguistic groups and 364 varieties. As they face a lack of representation in education, government, public services and media, they prese…
DiversityAxolotl: a Web Accessible Parallel Corpus for Spanish-Nahuatl
This paper describes the project called Axolotl which comprises a Spanish-Nahuatl parallel corpus and its search interface. Spanish and Nahuatl are distant languages spoken in the same country. Due to the scarcity of dig…
SentenceTowards an Open Source Finite-State Morphological Analyzer for Zacatlán-Ahuacatlán-Tepetzintla Nahuatl
Segmentation of nearly isotropic overlapped tracks in photomicrographs using successive erosions as watershed markers
The major challenges of automatic track counting are distinguishing tracks and material defects, identifying small tracks and defects of similar size, and detecting overlapping tracks. Here we address the latter issue us…