Weakly-Supervised Concept-based Adversarial Learning for Cross-lingual Word Embeddings
Distributed representations of words which map each word to a continuous vector have proven useful in capturing important linguistic information not only in a single language but also across different languages. Current unsupervised adversarial approaches show that it is possible to build a mapping matrix that align two sets of monolingual word embeddings together without high quality parallel data such as a dictionary or a sentence-aligned corpus. However, without post refinement, the performance of these methods' preliminary mapping is not good, leading to poor performance for typologically distant languages. In this paper, we propose a weakly-supervised adversarial training method to overcome this limitation, based on the intuition that mapping across languages is better done at the concept level than at the word level. We propose a concept-based adversarial training method which for most languages improves the performance of previous unsupervised adversarial methods, especially for typologically distant language pairs.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual Word EmbeddingsSentenceWord EmbeddingsSimilar Papers 제목 키워드 기반
Do We Really Need Fully Unsupervised Cross-Lingual Embeddings?
Recent efforts in cross-lingual word embedding (CLWE) learning have predominantly focused on fully unsupervised approaches that project monolingual embeddings into a shared cross-lingual space without any cross-lingual s…
Bilingual Lexicon InductionSelf-LearningUnsupervised Cross-Lingual Representation Learning
In this tutorial, we provide a comprehensive survey of the exciting recent work on cutting-edge weakly-supervised and unsupervised cross-lingual word representations. After providing a brief history of supervised cross-l…
Representation LearningStructured PredictionMulti-SimLex: A Large-Scale Evaluation of Multilingual and Crosslingual Lexical Semantic Similarity
We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering data sets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well a…
Representation LearningSemantic SimilaritySemantic Textual SimilarityWord EmbeddingsClassification-Based Self-Learning for Weakly Supervised Bilingual Lexicon Induction
Effective projection-based cross-lingual word embedding (CLWE) induction critically relies on the iterative self-learning procedure. It gradually expands the initial small seed dictionary to learn improved cross-lingual …
Bilingual Lexicon InductionClassificationGeneral ClassificationSelf-Learning+1