paper-with-me

Papers Multilingual Word Embeddings

“Multilingual Word Embeddings” 태그가 달린 논문 57편 · 필터 해제

How does a Multilingual LM Handle Multiple Languages?

2025-02-06 · Santhosh Kakarla, Gautama Shastry Bulusu Venkata, Aishwarya Gaddam

Multilingual language models have significantly advanced due to rapid progress in natural language processing. Models like BLOOM 1.7B, trained on diverse multilingual datasets, aim to bridge linguistic gaps. However, the…

Multilingual NLPMultilingual Word Embeddingsnamed-entity-recognitionNamed Entity Recognition+9

Multilingual Word Embeddings for Low-Resource Languages using Anchors and a Chain of Related Languages

2023-11-21 · Viktor Hangya, Silvia Severini, Radoslav Ralev, Alexander Fraser 외

Very low-resource languages, having only a few million tokens worth of data, are not well-supported by multilingual NLP approaches due to poor quality cross-lingual word representations. Recent work showed that good cros…

Bilingual Lexicon InductionMultilingual NLPMultilingual Word EmbeddingsWord Embeddings

OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining

2023-11-15 · Yihong Liu, Peiqin Lin, Mingyang Wang, Hinrich Schütze

Instead of pretraining multilingual language models from scratch, a more efficient method is to adapt existing pretrained language models (PLMs) to new languages via vocabulary extension and continued pretraining. Howeve…

Language ModellingMultilingual Word EmbeddingsWord Embeddings

Language Embeddings Sometimes Contain Typological Generalizations

2023-01-19 · Robert Östling, Murathan Kurfali

To what extent can neural network models learn generalizations about language structure, and how do we find out what they have learned? We explore these questions by training neural models for a range of natural language…

Multilingual Word EmbeddingsWord Embeddings

Improving Bilingual Lexicon Induction with Cross-Encoder Reranking

2022-10-30 · Yaoyiran Li, Fangyu Liu, Ivan Vulić, Anna Korhonen

Bilingual lexicon induction (BLI) with limited bilingual supervision is a crucial yet challenging task in multilingual NLP. Current state-of-the-art BLI methods rely on the induction of cross-lingual word embeddings (CLW…

Bilingual Lexicon InductionCross Encoder RerankingCross-Lingual Word EmbeddingsMachine Translation+9

Caveats of Measuring Semantic Change of Cognates and Borrowings using Multilingual Word Embeddings

2022-05-01 · LChange (ACL) 2022 5 · Clémentine Fourrier, Syrielle Montariol

Cognates and borrowings carry different aspects of etymological evolution. In this work, we study semantic change of such items using multilingual word embeddings, both static and contextualised. We underline caveats ide…

Multilingual Word EmbeddingsWord Embeddings

Improving Word Translation via Two-Stage Contrastive Learning

2022-03-15 · ACL 2022 5 · Yaoyiran Li, Fangyu Liu, Nigel Collier, Anna Korhonen 외

Word translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages. In this work, we propose a robust and effective two-stage contrastive learn…

Bilingual Lexicon InductionContrastive LearningCross-Lingual Word EmbeddingsMachine Translation+7

Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision

2022-02-16 · Priya Goyal, Quentin Duval, Isaac Seessel, Mathilde Caron 외

Discriminative self-supervised learning allows training models on any random group of internet images, and possibly recover salient information that helps differentiate between the images. Applied to ImageNet, this leads…

Action ClassificationAction RecognitionCopy DetectionDomain Generalization+13

Improving Word Translation via Two-Stage Contrastive Learning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Word translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages. In this work, we propose a robust and effective two-stage contrastive learn…

Bilingual Lexicon InductionContrastive LearningCross-Lingual Word EmbeddingsMachine Translation+8

Zero-Shot Cross-Lingual Transfer is a Hard Baseline to Beat in German Fine-Grained Entity Typing

2021-11-01 · EMNLP (insights) 2021 11 · Sabine Weber, Mark Steedman

The training of NLP models often requires large amounts of labelled training data, which makes it difficult to expand existing models to new languages. While zero-shot cross-lingual transfer relies on multilingual word e…

Cross-Lingual TransferEntity TypingMultilingual Word Embeddingsnamed-entity-recognition+4

Multilingual Dependency Parsing for Low-Resource African Languages: Case Studies on Bambara, Wolof, and Yoruba

2021-08-01 · ACL (IWPT) 2021 8 · Cheikh M. Bamba Dione

This paper describes a methodology for syntactic knowledge transfer between high-resource languages to extremely low-resource languages. The methodology consists in leveraging multilingual BERT self-attention model pretr…

Dependency ParsingMultilingual Word EmbeddingsTransfer LearningWord Embeddings

Debiasing Multilingual Word Embeddings: A Case Study of Three Indian Languages

2021-07-21 · Srijan Bansal, Vishal Garimella, Ayush Suhane, Animesh Mukherjee

In this paper, we advance the current state-of-the-art method for debiasing monolingual word embeddings so as to generalize well in a multilingual setting. We consider different methods to quantify bias and different deb…

Multilingual Word EmbeddingsWord Embeddings

A Primer on Pretrained Multilingual Language Models

2021-07-01 · Sumanth Doddapaneni, Gowtham Ramesh, Mitesh M. Khapra, Anoop Kunchukuttan 외

Multilingual Language Models (\MLLMs) such as mBERT, XLM, XLM-R, \textit{etc.} have emerged as a viable option for bringing the power of pretraining to a large number of languages. Given their success in zero-shot transf…

Joint Multilingual Sentence RepresentationsMultilingual text classificationMultilingual Word EmbeddingsPretrained Multilingual Language Models+3

Bootstrapping Multilingual AMR with Contextual Word Alignments

2021-02-03 · EACL 2021 2 · Janaki Sheth, Young-suk Lee, Ramon Fernandez Astudillo, Tahira Naseem 외

We develop high performance multilingualAbstract Meaning Representation (AMR) sys-tems by projecting English AMR annotationsto other languages with weak supervision. Weachieve this goal by bootstrapping transformer-based…

Multilingual Word EmbeddingsWord AlignmentWord EmbeddingsXLM-R

SHIKEBLCU at SemEval-2020 Task 2: An External Knowledge-enhanced Matrix for Multilingual and Cross-Lingual Lexical Entailment

2020-12-01 · SEMEVAL 2020 · Shike Wang, Yuchen Fan, Xiangying Luo, Dong Yu

Lexical entailment recognition plays an important role in tasks like Question Answering and Machine Translation. As important branches of lexical entailment, predicting multilingual and cross-lingual lexical entailment (…

Lexical EntailmentMachine TranslationMultilingual Word EmbeddingsQuestion Answering+3

Alignment-free Cross-lingual Semantic Role Labeling

2020-11-01 · EMNLP 2020 11 · Rui Cai, Mirella Lapata

Cross-lingual semantic role labeling (SRL) aims at leveraging resources in a source language to minimize the effort required to construct annotations or models for a new target language. Recent approaches rely on word al…

Machine TranslationMultilingual Word EmbeddingsSemantic Role LabelingTranslation+1

CS-Embed at SemEval-2020 Task 9: The effectiveness of code-switched word embeddings for sentiment analysis

2020-06-08 · SEMEVAL 2020 · Frances Adriana Laureano De Leon, Florimond Guéniat, Harish Tayyar Madabushi

The growing popularity and applications of sentiment analysis of social media posts has naturally led to sentiment analysis of posts written in multiple languages, a practice known as code-switching. While recent researc…

Multilingual Word EmbeddingsSentiment AnalysisWord Embeddings

Line-a-line: A Tool for Annotating Word-Alignments

2020-05-01 · LREC 2020 5 · Maria Skeppstedt, Magnus Ahltorp, Gunnar Eriksson, Rickard Domeij

We here describe line-a-line, a web-based tool for manual annotation of word-alignments in sentence-aligned parallel corpora. The graphical user interface, which builds on a design template from the Jigsaw system for inv…

Multilingual Word EmbeddingsSentenceWord AlignmentWord Embeddings

SimAlign: High Quality Word Alignments without Parallel Training Data using Static and Contextualized Embeddings

2020-04-18 · Findings of the Association for Computational Linguistics 2020 · Masoud Jalili Sabet, Philipp Dufter, François Yvon, Hinrich Schütze

Word alignments are useful for tasks like statistical and neural machine translation (NMT) and cross-lingual annotation projection. Statistical word aligners perform well, as do methods that extract alignments jointly wi…

Machine TranslationMultilingual Word EmbeddingsNMTTranslation+2

A Comparison of Architectures and Pretraining Methods for Contextualized Multilingual Word Embeddings

2019-12-15 · Niels van der Heijden, Samira Abnar, Ekaterina Shutova

The lack of annotated data in many languages is a well-known challenge within the field of multilingual natural language processing (NLP). Therefore, many recent studies focus on zero-shot transfer learning and joint tra…

Multilingual Word Embeddingsnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+7
1–20 / 57 다음 →