Predictive Text for Agglutinative and Polysynthetic Languages
This paper presents a set of experiments in the area of morphological modelling and prediction. We test whether morphological segmentation can compete against statistical segmentation in the tasks of language modelling and predictive text entry for two under-resourced and indigenous languages, K’iche’ and Chukchi. We use different segmentation methods — both statistical and morphological — to make datasets that are used to train models of different types: single-way segmented, which are trained using data from one segmenter; two-way segmented, which are trained using concatenated data from two segmenters; and finetuned, which are trained on two datasets from different segmenters. We compute word and character level perplexities and find that single-way segmented models trained on morphologically segmented data show the highest performance. Finally, we evaluate the language models on the task of predictive text entry using gold standard data and measure the average number of clicks per character and keystroke savings rate. We find that the models trained on morphologically segmented data show better scores, although with substantial room for improvement. At last, we propose the usage of morphological segmentation in order to improve the end-user experience while using predictive text and we plan on testing this assumption by doing end-user evaluation.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModellingSegmentationSimilar Papers 제목 키워드 기반
Predictive text for agglutinative and polysynthetic languages
This paper presents a set of experiments in the area of morphological modelling and predictioning. We examine the tasks of segmentation and predictive text entry for two under-resourced and indigenous languages, K'iche'…
Language ModellingSegmentationAutomatic Transcription Challenges for Inuktitut, a Low-Resource Polysynthetic Language
We introduce the first attempt at automatic speech recognition (ASR) in Inuktitut, as a representative for polysynthetic, low-resource languages, like many of the 900 Indigenous languages spoken in the Americas. As most …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionCorpora deduplication or duplication in Natural Language Processing of few resourced languages ? A case of study: The Mexico's Nahuatl
In this article, we seek to answer the following question: could data duplication be useful in Natural Language Processing (NLP) for languages with limited computational resources? In this type of languages (or $π$-langu…
Semantic SimilarityLost in Translation: Analysis of Information Loss During Machine Translation Between Polysynthetic and Fusional Languages
Machine translation from polysynthetic to fusional languages is a challenging task, which gets further complicated by the limited amount of parallel text available. Thus, translation performance is far from the state of …
Machine TranslationTranslationUnsupervised Morphological Segmentation for Low-Resource Polysynthetic Languages
Polysynthetic languages pose a challenge for morphological analysis due to the root-morpheme complexity and to the word class {``}squish{''}. In addition, many of these polysynthetic languages are low-resource. We propos…
Morphological Analysis