Embedding Learning Through Multilingual Concept Induction
We present a new method for estimating vector space representations of words: embedding learning by concept induction. We test this method on a highly parallel corpus and learn semantic representations of words in 1259 different languages in a single common space. An extensive experimental evaluation on crosslingual word similarity and sentiment analysis indicates that concept-based multilingual embedding learning performs better than previous approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
Sentiment AnalysisWord SimilaritySimilar Papers 제목 키워드 기반
A Multilingual BPE Embedding Space for Universal Sentiment Lexicon Induction
We present a new method for sentiment lexicon induction that is designed to be applicable to the entire range of typological diversity of the world{'}s languages. We evaluate our method on Parallel Bible Corpus+ (PBC+), …
DiversityDomain AdaptationLearning Unsupervised Multilingual Word Embeddings with Incremental Multilingual Hubs
Recent research has discovered that a shared bilingual word embedding space can be induced by projecting monolingual word embedding spaces from two languages using a self-learning paradigm without any bilingual supervisi…
Bilingual Lexicon InductionCross-Lingual Word EmbeddingsDependency ParsingDocument Classification+3A Simple Approach to Learning Unsupervised Multilingual Embeddings
Recent progress on unsupervised learning of cross-lingual embeddings in bilingual setting has given impetus to learning a shared embedding space for several languages without any supervision. A popular framework to solve…
Bilingual Lexicon InductionDependency ParsingDocument ClassificationWord Alignment+1Multilingual Sentence Transformer as A Multilingual Word Aligner
Multilingual pretrained language models (mPLMs) have shown their effectiveness in multilingual word alignment induction. However, these methods usually start from mBERT or XLM-R. In this paper, we investigate whether mul…
SentenceWord AlignmentXLM-RShould All Cross-Lingual Embeddings Speak English?
Most of recent work in cross-lingual word embeddings is severely Anglocentric. The vast majority of lexicon induction evaluation dictionaries are between English and another language, and the English embedding space is s…
AllCross-Lingual Word EmbeddingsWord Embeddings