Interactive Refinement of Cross-Lingual Word Embeddings
Cross-lingual word embeddings transfer knowledge between languages: models trained on high-resource languages can predict in low-resource languages. We introduce CLIME, an interactive system to quickly refine cross-lingual word embeddings for a given classification problem. First, CLIME ranks words by their salience to the downstream task. Then, users mark similarity between keywords and their nearest neighbors in the embedding space. Finally, CLIME updates the embeddings using the annotations. We evaluate CLIME on identifying health-related text in four low-resource languages: Ilocano, Sinhalese, Tigrinya, and Uyghur. Embeddings refined by CLIME capture more nuanced word semantics and have higher test accuracy than the original embeddings. CLIME often improves accuracy faster than an active learning baseline and can be easily combined with active learning to improve results.
Code (1)
Tasks
Active LearningCross-Lingual Word EmbeddingsGeneral ClassificationText ClassificationWord EmbeddingsMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Multi-Stage Framework with Refinement Based Point Set Registration for Unsupervised Bi-Lingual Word Alignment
Cross-lingual alignment of word embeddings are important in knowledge transfer across languages, for improving machine translation and other multi-lingual applications. Current unsupervised approaches relying on learning…
Machine TranslationTransfer LearningTranslationWord Alignment+2Refinement of Unsupervised Cross-Lingual Word Embeddings
Cross-lingual word embeddings aim to bridge the gap between high-resource and low-resource languages by allowing to learn multilingual word representations even without using any direct bilingual signal. The lion's share…
Bilingual Lexicon InductionCross-Lingual Word EmbeddingsWord EmbeddingsMulti-Stage Framework with Refinement based Point Set Registration for Unsupervised Bi-Lingual Word Alignment
Cross-lingual alignment of word embeddings play an important role in knowledge transfer across languages, for improving machine translation and other multi-lingual applications. Current unsupervised approaches rely on le…
Machine TranslationSentenceTransfer LearningTranslation+3Unsupervised Word Translation with Adversarial Autoencoder
Crosslingual word embeddings learned from monolingual embeddings have a crucial role in many downstream tasks, ranging from machine translation to transfer learning. Adversarial training has shown impressive success in l…
Machine TranslationTransfer LearningTranslationWord Embeddings+1Unsupervised Word Translation Pairing using Refinement based Point Set Registration
Cross-lingual alignment of word embeddings play an important role in knowledge transfer across languages, for improving machine translation and other multi-lingual applications. Current unsupervised approaches rely on si…
Machine TranslationTransfer LearningTranslationWord Embeddings+1