paper-with-me

홈 › Papers

Massive vs. Curated Word Embeddings for Low-Resourced Languages. The Case of Yorùbá and Twi

2019-12-05 · Jesujoba O. Alabi, Kwabena Amponsah-Kaakyire, David I. Adelani, Cristina España-Bonet

The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and automatic creation of resources for multiple languages. The evaluation of such resources is usually done for the high-resourced languages, where one has a smorgasbord of tasks and test sets to evaluate on. For low-resourced languages, the evaluation is more difficult and normally ignored, with the hope that the impressive capability of deep learning architectures to learn (multilingual) representations in the high-resourced setting holds in the low-resourced setting too. In this paper we focus on two African languages, Yor\ub\'a and Twi, and compare the word embeddings obtained in this way, with word embeddings obtained from curated corpora and a language-dependent processing. We analyse the noise in the publicly available corpora, collect high quality and noisy data for the two languages and quantify the improvements that depend not only on the amount of data but on the quality too. We also use different architectures that learn word representations both from surface forms and characters to further exploit all the available information which showed to be important for these languages. For the evaluation, we manually translate the wordsim-353 word pairs dataset from English into Yor\ub\'a and Twi. As output of the work, we provide corpora, embeddings and the test suits for both languages.

📄 PDF Abstract BibTeX arXiv:1912.02481

Code (1)

ajesujoba/YorubaTwi-Embedding 공식 구현

Tasks

Word Embeddings

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Massive vs. Curated Embeddings for Low-Resourced Languages: the Case of Yor\`ub\'a and Twi

2020-05-01 · LREC 2020 5 · Jesujoba Alabi, Kwabena Amponsah-Kaakyire, David Adelani, Cristina Espa{\~n}a-Bonet

The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Word Embeddings

When Word Embeddings Become Endangered

2021-03-24 · Khalid Alnajjar

Big languages such as English and Finnish have many natural language processing (NLP) resources and models, but this is not the case for low-resourced and endangered languages as such resources are so scarce despite the …

Cross-Lingual Word EmbeddingsSentiment AnalysisTranslationWord Embeddings

Supervised and Nonlinear Alignment of Two Embedding Spaces for Dictionary Induction in Low Resourced Languages

2019-11-01 · IJCNLP 2019 11 · Masud Moshtaghi

Enabling cross-lingual NLP tasks by leveraging multilingual word embedding has recently attracted much attention. An important motivation is to support lower resourced languages, however, most efforts focus on demonstrat…

Multilingual acoustic word embedding models for processing zero-resource languages

2020-02-06 · Herman Kamper, Yevgen Matusevych, Sharon Goldwater

Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments. In settings where unlabelled speech is the only available resource, such embeddings can be used in "zero-resource" speech…

Transfer LearningWord Embeddings

Multilingual acoustic word embeddings for zero-resource languages

2024-01-19 · Christiaan Jacobs

This research addresses the challenge of developing speech applications for zero-resource languages that lack labelled data. It specifically uses acoustic word embedding (AWE) -- fixed-dimensional representations of vari…

Hate Speech DetectionKeyword SpottingWord Embeddings