Unsupervised acquisition of concatenative morphology
Among the linguistic resources formalizing a language, morphological rules are among those that can be achieved in a reasonable time. Nevertheless, since the construction of such resource can require linguistic expertise, morphological rules are still lacking for many languages. The automatized acquisition of morphology is thus an open topic of interest within the NLP field. We present an approach that allows to automatically compute, from raw corpora, a data-representative description of the concatenative mechanisms of a morphology. Our approach takes advantage of phenomena that are observable for all languages using morphological inflection and derivation but are more easy to exploit when dealing with concatenative mechanisms. Since it has been developed toward the objective of being used on as many languages as possible, applying this approach to a varied set of languages needs very few expert work. The results obtained for our first participation in the 2010 edition of MorphoChallenge have confirmed both the practical interest and the potential of the method.
Code (0)
등록된 구현이 없습니다.
Tasks
Morphological InflectionSimilar Papers 제목 키워드 기반
Learning non-concatenative morphology
Adaptor Grammars for Learning Non-Concatenative Morphology
How Suitable Are Subword Segmentation Strategies for Translating Non-Concatenative Morphology?
Data-driven subword segmentation has become the default strategy for open-vocabulary machine translation and other NLP tasks, but may not be sufficiently generic for optimal learning of non-concatenative morphology. We d…
Machine TranslationSegmentationTranslationSplintering Nonconcatenative Languages for Better Tokenization
Common subword tokenization algorithms like BPE and UnigramLM assume that text can be split into meaningful units by concatenative measures alone. This is not true for languages such as Hebrew and Arabic, where morpholog…
Constrained Sequence-to-sequence Semitic Root Extraction for Enriching Word Embeddings
In this paper, we tackle the problem of {``}root extraction{''} from words in the Semitic language family. A challenge in applying natural language processing techniques to these languages is the data sparsity problem th…
Language ModelingLanguage ModellingWord EmbeddingsWord Similarity