Evaluating Word Embeddings on Low-Resource Languages
The analogy task introduced by Mikolov et al. (2013) has become the standard metric for tuning the hyperparameters of word embedding models. In this paper, however, we argue that the analogy task is unsuitable for low-resource languages for two reasons: (1) it requires that word embeddings be trained on large amounts of text, and (2) analogies may not be well-defined in some low-resource settings. We solve these problems by introducing the OddOneOut and Topk tasks, which are specifically designed for model selection in the low-resource setting. We use these metrics to successfully tune hyperparameters for a low-resource emoji embedding task and word embeddings on 16 extinct languages. The largest of these languages (Ancient Hebrew) has a 41 million token dataset, and the smallest (Old Gujarati) has only a 1813 token dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Model SelectionWord EmbeddingsSimilar Papers 제목 키워드 기반
Evaluating Sub-word Embeddings in Cross-lingual Models
Cross-lingual word embeddings create a shared space for embeddings in two languages, and enable knowledge to be transferred between languages for tasks such as bilingual lexicon induction. One problem, however, is out-of…
Bilingual Lexicon InductionCross-Lingual Word EmbeddingsWord EmbeddingsEvaluating a Joint Training Approach for Learning Cross-lingual Embeddings with Sub-word Information without Parallel Corpora on Lower-resource Languages
Cross-lingual word embeddings provide a way for information to be transferred between languages. In this paper we evaluate an extension of a joint training approach to learning cross-lingual embeddings that incorporates …
Bilingual Lexicon InductionCross-Lingual Word EmbeddingsWord EmbeddingsGlobalTrait: Personality Alignment of Multilingual Word Embeddings
We propose a multilingual model to recognize Big Five Personality traits from text data in four different languages: English, Spanish, Dutch and Italian. Our analysis shows that words having a similar semantic meaning in…
Multilingual Word EmbeddingsPersonality AlignmentWord EmbeddingsWhen Word Embeddings Become Endangered
Big languages such as English and Finnish have many natural language processing (NLP) resources and models, but this is not the case for low-resourced and endangered languages as such resources are so scarce despite the …
Cross-Lingual Word EmbeddingsSentiment AnalysisTranslationWord EmbeddingsCoSimLex: A Resource for Evaluating Graded Word Similarity in Context
State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic eva…
Word EmbeddingsWord Sense DisambiguationWord Similarity