Morphology-rich Alphasyllabary Embeddings
Word embeddings have been successfully trained in many languages. However, both intrinsic and extrinsic metrics are variable across languages, especially for languages that depart significantly from English in morphology and orthography. This study focuses on building a word embedding model suitable for the Semitic language of Amharic (Ethiopia), which is both morphologically rich and written as an alphasyllabary (abugida) rather than an alphabet. We compare embeddings from tailored neural models, simple pre-processing steps, off-the-shelf baselines, and parallel tasks on a better-resourced Semitic language {--} Arabic. Experiments show our model{'}s performance on word analogy tasks, illustrating the divergent objectives of morphological vs. semantic analogies.
Code (0)
등록된 구현이 없습니다.
Tasks
Word EmbeddingsSimilar Papers 제목 키워드 기반
Telugu OCR Framework using Deep Learning
In this paper, we address the task of Optical Character Recognition(OCR) for the Telugu script. We present an end-to-end framework that segments the text image, classifies the characters and extracts lines using a langua…
Deep LearningGeneral ClassificationLanguage ModelingLanguage Modelling+2Evaluation of Morphological Embeddings for the Russian Language
A number of morphology-based word embedding models were introduced in recent years. However, their evaluation was mostly limited to English, which is known to be a morphologically simple language. In this paper, we explo…
ChunkingNERPOSPOS Tagging+1Morphology-Aware Meta-Embeddings for Tamil
In this work, we explore generating morphologically enhanced word embeddings for Tamil, a highly agglutinative South Indian language with rich morphology that remains low-resource with regards to NLP tasks. We present he…
Word EmbeddingsChinese Embedding via Stroke and Glyph Information: A Dual-channel View
Recent studies have consistently given positive hints that morphology is helpful in enriching word embeddings. In this paper, we argue that Chinese word embeddings can be substantially enriched by the morphological infor…
Word EmbeddingsWord SimilarityConstrained Sequence-to-sequence Semitic Root Extraction for Enriching Word Embeddings
In this paper, we tackle the problem of {``}root extraction{''} from words in the Semitic language family. A challenge in applying natural language processing techniques to these languages is the data sparsity problem th…
Language ModelingLanguage ModellingWord EmbeddingsWord Similarity