Papers Pretrained Multilingual Language Models
“Pretrained Multilingual Language Models” 태그가 달린 논문 26편 · 필터 해제
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
We present OpenNER 1.0, a standardized collection of openly available named entity recognition (NER) datasets. OpenNER contains 34 datasets spanning 51 languages, annotated in varying named entity ontologies. We correct …
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1Discovering Low-rank Subspaces for Language-agnostic Multilingual Representations
Large pretrained multilingual language models (ML-LMs) have shown remarkable capabilities of zero-shot cross-lingual transfer, without direct cross-lingual supervision. While these results are promising, follow-up works …
Cross-Lingual TransferPretrained Multilingual Language ModelsRetrievalSentence+2Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation
Zero-shot cross-lingual knowledge transfer enables the multilingual pretrained language model (mPLM), finetuned on a task in one language, make predictions for this task in other languages. While being broadly studied fo…
Language ModelingLanguage ModellingNatural Language UnderstandingPretrained Multilingual Language Models+1Exploring the Maze of Multilingual Modeling
Multilingual language models have gained significant attention in recent years, enabling the development of applications that meet diverse linguistic contexts. In this paper, we present a comprehensive evaluation of thre…
Language ModellingModel SelectionPretrained Multilingual Language Modelstext-classification+3Improving Non-autoregressive Translation Quality with Pretrained Language Model, Embedding Distillation and Upsampling Strategy for CTC
Non-autoregressive approaches aim to improve the inference speed of translation models, particularly those that generate output in a one-pass forward manner. However, these approaches often suffer from a significant drop…
Language ModelingLanguage ModellingPretrained Multilingual Language ModelsTranslationImproving Bilingual Lexicon Induction with Cross-Encoder Reranking
Bilingual lexicon induction (BLI) with limited bilingual supervision is a crucial yet challenging task in multilingual NLP. Current state-of-the-art BLI methods rely on the induction of cross-lingual word embeddings (CLW…
Bilingual Lexicon InductionCross Encoder RerankingCross-Lingual Word EmbeddingsMachine Translation+9Language Agnostic Multilingual Information Retrieval with Contrastive Learning
Multilingual information retrieval (IR) is challenging since annotated training data is costly to obtain in many languages. We present an effective method to train multilingual IR systems when only English IR training da…
Contrastive LearningCross-Lingual TransferInformation RetrievalPretrained Multilingual Language Models+3Are Pretrained Multilingual Models Equally Fair Across Languages?
Pretrained multilingual language models can help bridge the digital language divide, enabling high-quality NLP models for lower resourced languages. Studies of multilingual models have so far focused on performance, cons…
Cloze TestFairnessPretrained Multilingual Language ModelsXLM-RRobustification of Multilingual Language Models to Real-world Noise in Crosslingual Zero-shot Settings with Robust Contrastive Pretraining
Advances in neural modeling have achieved state-of-the-art (SOTA) results on public natural language processing (NLP) benchmarks, at times surpassing human performance. However, there is a gap between public benchmarks a…
Data AugmentationPretrained Multilingual Language ModelsSentenceTo Adapt or to Fine-tune: A Case Study on Abstractive Summarization
Recent advances in the field of abstractive summarization leverage pre-trained language models rather than train a model from scratch. However, such models are sluggish to train and accompanied by a massive overhead. Res…
Abstractive Text SummarizationLanguage ModelingLanguage ModellingPretrained Multilingual Language ModelsProbing Cross-Lingual Lexical Knowledge from Multilingual Sentence Encoders
Pretrained multilingual language models (LMs) can be successfully transformed into multilingual sentence encoders (SEs; e.g., LaBSE, xMPNet) via additional fine-tuning or model distillation with parallel data. However, i…
Contrastive LearningCross-Lingual Entity LinkingEntity LinkingPretrained Multilingual Language Models+4Team ÚFAL at CMCL 2022 Shared Task: Figuring out the correct recipe for predicting Eye-Tracking features using Pretrained Language Models
Eye-Tracking data is a very useful source of information to study cognition and especially language comprehension in humans. In this paper, we describe our systems for the CMCL 2022 shared task on predicting eye-tracking…
Pretrained Multilingual Language ModelsImproving Word Translation via Two-Stage Contrastive Learning
Word translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages. In this work, we propose a robust and effective two-stage contrastive learn…
Bilingual Lexicon InductionContrastive LearningCross-Lingual Word EmbeddingsMachine Translation+7Out of Thin Air: Is Zero-Shot Cross-Lingual Keyword Detection Better Than Unsupervised?
Keyword extraction is the task of retrieving words that are essential to the content of a given document. Researchers proposed various approaches to tackle this problem. At the top-most level, approaches are divided into…
Keyword ExtractionPretrained Multilingual Language ModelsAre Pretrained Multilingual Models Equally Fair Across Languages?
Pretrained multilingual language models can help bridge the digital language divide, enabling high-quality NLP models for lower-resourced languages. Studies of multilingual models have so far focused on performance, cons…
Cloze TestFairnessPretrained Multilingual Language ModelsXLM-RInvestigating Math Word Problems using Pretrained Multilingual Language Models
In this paper, we revisit math word problems~(MWPs) from the {\em cross-lingual} and {\em multilingual} perspective.We construct our MWP solvers over pretrained multilingual language models using the sequence-to-sequence…
Machine TranslationMathPretrained Multilingual Language ModelsTranslationMulti-Source Cross-Lingual Constituency Parsing
Pretrained multilingual language models have become a key part of cross-lingual transfer for many natural language processing tasks, even those without bilingual information. This work further investigates the cross-ling…
Constituency ParsingCross-Lingual TransferDiversityPretrained Multilingual Language ModelsImproving Word Translation via Two-Stage Contrastive Learning
Word translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages. In this work, we propose a robust and effective two-stage contrastive learn…
Bilingual Lexicon InductionContrastive LearningCross-Lingual Word EmbeddingsMachine Translation+8Small Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages
Pretrained multilingual language models have been shown to work well on many languages for a variety of downstream NLP tasks. However, these models are known to require a lot of training data. This consequently leaves ou…
Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+4Transliteration: A Simple Technique For Improving Multilingual Language Modeling
While impressive performance in natural language processing tasks has been achieved for many languages by transfer learning from large pretrained multilingual language models, it is limited by the unavailability of large…
Language ModelingLanguage ModellingMultiple-choiceMultiple Choice Question Answering (MCQA)+5