Morphological Segmentation for Low Resource Languages
This paper describes a new morphology resource created by Linguistic Data Consortium and the University of Pennsylvania for the DARPA LORELEI Program. The data consists of approximately 2000 tokens annotated for morphological segmentation in each of 9 low resource languages, along with root information for 7 of the languages. The languages annotated show a broad diversity of typological features. A minimal annotation scheme for segmentation was developed such that it could capture the patterns of a wide range of languages and also be performed reliably by non-linguist annotators. The basic annotation guidelines were designed to be language-independent, but included language-specific morphological paradigms and other specifications. The resulting annotated corpus is designed to support and stimulate the development of unsupervised morphological segmenters and analyzers by providing a gold standard for their evaluation on a more typologically diverse set of languages than has previously been available. By providing root annotation, this corpus is also a step toward supporting research in identifying richer morphological structures than simple morpheme boundaries.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversitySegmentationSimilar Papers 제목 키워드 기반
MorphAGram, Evaluation and Framework for Unsupervised Morphological Segmentation
Computational morphological segmentation has been an active research topic for decades as it is beneficial for many natural language processing tasks. With the high cost of manually labeling data for morphology and the i…
SegmentationUnsupervised Morphological Segmentation for Low-Resource Polysynthetic Languages
Polysynthetic languages pose a challenge for morphological analysis due to the root-morpheme complexity and to the word class {``}squish{''}. In addition, many of these polysynthetic languages are low-resource. We propos…
Morphological AnalysisFortification of Neural Morphological Segmentation Models for Polysynthetic Minimal-Resource Languages
Morphological segmentation for polysynthetic languages is challenging, because a word may consist of many individual morphemes and training data can be extremely scarce. Since neural sequence-to-sequence (seq2seq) models…
Cross-Lingual TransferData AugmentationSegmentationTowards a First Automatic Unsupervised Morphological Segmentation for Inuinnaqtun
Low-resource polysynthetic languages pose many challenges in NLP tasks, such as morphological analysis and Machine Translation, due to available resources and tools, and the morphologically complex languages. This resear…
Machine TranslationMorphological AnalysisTranslationAutomatically Tailoring Unsupervised Morphological Segmentation to the Language
Morphological segmentation is beneficial for several natural language processing tasks dealing with large vocabularies. Unsupervised methods for morphological segmentation are essential for handling a diverse set of lang…
Machine TranslationSegmentationSpeech Recognition