paper-with-me

홈 › Papers

An Unsupervised Approach for Mapping between Vector Spaces

2017-11-15 · Syed Sarfaraz Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava, Madan Gopal Jhawar, Manish Shrivastava

We present a language independent, unsupervised approach for transforming word embeddings from source language to target language using a transformation matrix. Our model handles the problem of data scarcity which is faced by many languages in the world and yields improved word embeddings for words in the target language by relying on transformed embeddings of words of the source language. We initially evaluate our approach via word similarity tasks on a similar language pair - Hindi as source and Urdu as the target language, while we also evaluate our method on French and German as target languages and English as source language. Our approach improves the current state of the art results - by 13% for French and 19% for German. For Urdu, we saw an increment of 16% over our initial baseline score. We further explore the prospects of our approach by applying it on multiple models of the same language and transferring words between the two models, thus solving the problem of missing words in a model. We evaluate this on word similarity and word analogy tasks.

📄 PDF Abstract BibTeX arXiv:1711.05680

Code (0)

등록된 구현이 없습니다.

Tasks

Word EmbeddingsWord Similarity

Similar Papers 제목 키워드 기반

Cross-topic distributional semantic representations via unsupervised mappings

2019-04-11 · NAACL 2019 6 · Eleftheria Briakou, Nikos Athanasiou, Alexandros Potamianos

In traditional Distributional Semantic Models (DSMs) the multiple senses of a polysemous word are conflated into a single vector space representation. In this work, we propose a DSM that learns multiple distributional re…

Word Similarity

Bootstrapping Unsupervised Bilingual Lexicon Induction

2017-04-01 · EACL 2017 4 · Bradley Hauer, Garrett Nicolai, Grzegorz Kondrak

The task of unsupervised lexicon induction is to find translation pairs across monolingual corpora. We develop a novel method that creates seed lexicons by identifying cognates in the vocabularies of related languages on…

Bilingual Lexicon InductionSemantic Textual SimilarityTranslation

Learning Unsupervised Word Mapping by Maximizing Mean Discrepancy

2018-11-01 · Pengcheng Yang, Fuli Luo, Shuangzhi Wu, Jingjing Xu 외

Cross-lingual word embeddings aim to capture common linguistic regularities of different languages, which benefit various downstream tasks ranging from machine translation to transfer learning. Recently, it has been show…

Cross-Lingual Word EmbeddingsDensity EstimationMachine TranslationTransfer Learning+2

Unsupervised Word Mapping Using Structural Similarities in Monolingual Embeddings

2017-12-19 · TACL 2018 1 · Hanan Aldarmaki, Mahesh Mohan, Mona Diab

Most existing methods for automatic bilingual dictionary induction rely on prior alignments between the source and target languages, such as parallel corpora or seed dictionaries. For many language pairs, such supervised…

Word Embeddings

Data Augmentation with Unsupervised Machine Translation Improves the Structural Similarity of Cross-lingual Word Embeddings

2020-05-30 · ACL 2021 5 · Sosuke Nishikawa, Ryokan Ri, Yoshimasa Tsuruoka

Unsupervised cross-lingual word embedding (CLWE) methods learn a linear transformation matrix that maps two monolingual embedding spaces that are separately trained with monolingual corpora. This method relies on the ass…

Cross-Lingual Word EmbeddingsData AugmentationMachine TranslationTranslation+2