Papers Multilingual Word Embeddings
“Multilingual Word Embeddings” 태그가 달린 논문 57편 · 필터 해제
How does a Multilingual LM Handle Multiple Languages?
Multilingual language models have significantly advanced due to rapid progress in natural language processing. Models like BLOOM 1.7B, trained on diverse multilingual datasets, aim to bridge linguistic gaps. However, the…
Multilingual NLPMultilingual Word Embeddingsnamed-entity-recognitionNamed Entity Recognition+9Multilingual Word Embeddings for Low-Resource Languages using Anchors and a Chain of Related Languages
Very low-resource languages, having only a few million tokens worth of data, are not well-supported by multilingual NLP approaches due to poor quality cross-lingual word representations. Recent work showed that good cros…
Bilingual Lexicon InductionMultilingual NLPMultilingual Word EmbeddingsWord EmbeddingsOFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
Instead of pretraining multilingual language models from scratch, a more efficient method is to adapt existing pretrained language models (PLMs) to new languages via vocabulary extension and continued pretraining. Howeve…
Language ModellingMultilingual Word EmbeddingsWord EmbeddingsLanguage Embeddings Sometimes Contain Typological Generalizations
To what extent can neural network models learn generalizations about language structure, and how do we find out what they have learned? We explore these questions by training neural models for a range of natural language…
Multilingual Word EmbeddingsWord EmbeddingsImproving Bilingual Lexicon Induction with Cross-Encoder Reranking
Bilingual lexicon induction (BLI) with limited bilingual supervision is a crucial yet challenging task in multilingual NLP. Current state-of-the-art BLI methods rely on the induction of cross-lingual word embeddings (CLW…
Bilingual Lexicon InductionCross Encoder RerankingCross-Lingual Word EmbeddingsMachine Translation+9Caveats of Measuring Semantic Change of Cognates and Borrowings using Multilingual Word Embeddings
Cognates and borrowings carry different aspects of etymological evolution. In this work, we study semantic change of such items using multilingual word embeddings, both static and contextualised. We underline caveats ide…
Multilingual Word EmbeddingsWord EmbeddingsImproving Word Translation via Two-Stage Contrastive Learning
Word translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages. In this work, we propose a robust and effective two-stage contrastive learn…
Bilingual Lexicon InductionContrastive LearningCross-Lingual Word EmbeddingsMachine Translation+7Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision
Discriminative self-supervised learning allows training models on any random group of internet images, and possibly recover salient information that helps differentiate between the images. Applied to ImageNet, this leads…
Action ClassificationAction RecognitionCopy DetectionDomain Generalization+13Improving Word Translation via Two-Stage Contrastive Learning
Word translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages. In this work, we propose a robust and effective two-stage contrastive learn…
Bilingual Lexicon InductionContrastive LearningCross-Lingual Word EmbeddingsMachine Translation+8Zero-Shot Cross-Lingual Transfer is a Hard Baseline to Beat in German Fine-Grained Entity Typing
The training of NLP models often requires large amounts of labelled training data, which makes it difficult to expand existing models to new languages. While zero-shot cross-lingual transfer relies on multilingual word e…
Cross-Lingual TransferEntity TypingMultilingual Word Embeddingsnamed-entity-recognition+4Multilingual Dependency Parsing for Low-Resource African Languages: Case Studies on Bambara, Wolof, and Yoruba
This paper describes a methodology for syntactic knowledge transfer between high-resource languages to extremely low-resource languages. The methodology consists in leveraging multilingual BERT self-attention model pretr…
Dependency ParsingMultilingual Word EmbeddingsTransfer LearningWord EmbeddingsDebiasing Multilingual Word Embeddings: A Case Study of Three Indian Languages
In this paper, we advance the current state-of-the-art method for debiasing monolingual word embeddings so as to generalize well in a multilingual setting. We consider different methods to quantify bias and different deb…
Multilingual Word EmbeddingsWord EmbeddingsA Primer on Pretrained Multilingual Language Models
Multilingual Language Models (\MLLMs) such as mBERT, XLM, XLM-R, \textit{etc.} have emerged as a viable option for bringing the power of pretraining to a large number of languages. Given their success in zero-shot transf…
Joint Multilingual Sentence RepresentationsMultilingual text classificationMultilingual Word EmbeddingsPretrained Multilingual Language Models+3Bootstrapping Multilingual AMR with Contextual Word Alignments
We develop high performance multilingualAbstract Meaning Representation (AMR) sys-tems by projecting English AMR annotationsto other languages with weak supervision. Weachieve this goal by bootstrapping transformer-based…
Multilingual Word EmbeddingsWord AlignmentWord EmbeddingsXLM-RSHIKEBLCU at SemEval-2020 Task 2: An External Knowledge-enhanced Matrix for Multilingual and Cross-Lingual Lexical Entailment
Lexical entailment recognition plays an important role in tasks like Question Answering and Machine Translation. As important branches of lexical entailment, predicting multilingual and cross-lingual lexical entailment (…
Lexical EntailmentMachine TranslationMultilingual Word EmbeddingsQuestion Answering+3Alignment-free Cross-lingual Semantic Role Labeling
Cross-lingual semantic role labeling (SRL) aims at leveraging resources in a source language to minimize the effort required to construct annotations or models for a new target language. Recent approaches rely on word al…
Machine TranslationMultilingual Word EmbeddingsSemantic Role LabelingTranslation+1CS-Embed at SemEval-2020 Task 9: The effectiveness of code-switched word embeddings for sentiment analysis
The growing popularity and applications of sentiment analysis of social media posts has naturally led to sentiment analysis of posts written in multiple languages, a practice known as code-switching. While recent researc…
Multilingual Word EmbeddingsSentiment AnalysisWord EmbeddingsLine-a-line: A Tool for Annotating Word-Alignments
We here describe line-a-line, a web-based tool for manual annotation of word-alignments in sentence-aligned parallel corpora. The graphical user interface, which builds on a design template from the Jigsaw system for inv…
Multilingual Word EmbeddingsSentenceWord AlignmentWord EmbeddingsSimAlign: High Quality Word Alignments without Parallel Training Data using Static and Contextualized Embeddings
Word alignments are useful for tasks like statistical and neural machine translation (NMT) and cross-lingual annotation projection. Statistical word aligners perform well, as do methods that extract alignments jointly wi…
Machine TranslationMultilingual Word EmbeddingsNMTTranslation+2A Comparison of Architectures and Pretraining Methods for Contextualized Multilingual Word Embeddings
The lack of annotated data in many languages is a well-known challenge within the field of multilingual natural language processing (NLP). Therefore, many recent studies focus on zero-shot transfer learning and joint tra…
Multilingual Word Embeddingsnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+7