paper-with-me

Papers Pretrained Multilingual Language Models

“Pretrained Multilingual Language Models” 태그가 달린 논문 26편 · 필터 해제

OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages

2024-12-12 · Chester Palen-Michel, Maxwell Pickering, Maya Kruse, Jonne Sälevä 외

We present OpenNER 1.0, a standardized collection of openly available named entity recognition (NER) datasets. OpenNER contains 34 datasets spanning 51 languages, annotated in varying named entity ontologies. We correct …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Discovering Low-rank Subspaces for Language-agnostic Multilingual Representations

2024-01-11 · Zhihui Xie, Handong Zhao, Tong Yu, Shuai Li

Large pretrained multilingual language models (ML-LMs) have shown remarkable capabilities of zero-shot cross-lingual transfer, without direct cross-lingual supervision. While these results are promising, follow-up works …

Cross-Lingual TransferPretrained Multilingual Language ModelsRetrievalSentence+2

Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation

2023-10-15 · Nadezhda Chirkova, Sheng Liang, Vassilina Nikoulina

Zero-shot cross-lingual knowledge transfer enables the multilingual pretrained language model (mPLM), finetuned on a task in one language, make predictions for this task in other languages. While being broadly studied fo…

Language ModelingLanguage ModellingNatural Language UnderstandingPretrained Multilingual Language Models+1

Exploring the Maze of Multilingual Modeling

2023-10-09 · Sina Bagheri Nezhad, Ameeta Agrawal

Multilingual language models have gained significant attention in recent years, enabling the development of applications that meet diverse linguistic contexts. In this paper, we present a comprehensive evaluation of thre…

Language ModellingModel SelectionPretrained Multilingual Language Modelstext-classification+3

Improving Non-autoregressive Translation Quality with Pretrained Language Model, Embedding Distillation and Upsampling Strategy for CTC

2023-06-10 · Shen-sian Syu, Juncheng Xie, Hung-Yi Lee

Non-autoregressive approaches aim to improve the inference speed of translation models, particularly those that generate output in a one-pass forward manner. However, these approaches often suffer from a significant drop…

Language ModelingLanguage ModellingPretrained Multilingual Language ModelsTranslation

Improving Bilingual Lexicon Induction with Cross-Encoder Reranking

2022-10-30 · Yaoyiran Li, Fangyu Liu, Ivan Vulić, Anna Korhonen

Bilingual lexicon induction (BLI) with limited bilingual supervision is a crucial yet challenging task in multilingual NLP. Current state-of-the-art BLI methods rely on the induction of cross-lingual word embeddings (CLW…

Bilingual Lexicon InductionCross Encoder RerankingCross-Lingual Word EmbeddingsMachine Translation+9

Language Agnostic Multilingual Information Retrieval with Contrastive Learning

2022-10-12 · Xiyang Hu, Xinchi Chen, Peng Qi, Deguang Kong 외

Multilingual information retrieval (IR) is challenging since annotated training data is costly to obtain in many languages. We present an effective method to train multilingual IR systems when only English IR training da…

Contrastive LearningCross-Lingual TransferInformation RetrievalPretrained Multilingual Language Models+3

Are Pretrained Multilingual Models Equally Fair Across Languages?

2022-10-11 · COLING 2022 10 · Laura Cabello Piqueras, Anders Søgaard

Pretrained multilingual language models can help bridge the digital language divide, enabling high-quality NLP models for lower resourced languages. Studies of multilingual models have so far focused on performance, cons…

Cloze TestFairnessPretrained Multilingual Language ModelsXLM-R

Robustification of Multilingual Language Models to Real-world Noise in Crosslingual Zero-shot Settings with Robust Contrastive Pretraining

2022-10-10 · Asa Cooper Stickland, Sailik Sengupta, Jason Krone, Saab Mansour 외

Advances in neural modeling have achieved state-of-the-art (SOTA) results on public natural language processing (NLP) benchmarks, at times surpassing human performance. However, there is a gap between public benchmarks a…

Data AugmentationPretrained Multilingual Language ModelsSentence

To Adapt or to Fine-tune: A Case Study on Abstractive Summarization

2022-08-30 · CCL 2022 10 · Zheng Zhao, Pinzhen Chen

Recent advances in the field of abstractive summarization leverage pre-trained language models rather than train a model from scratch. However, such models are sluggish to train and accompanied by a massive overhead. Res…

Abstractive Text SummarizationLanguage ModelingLanguage ModellingPretrained Multilingual Language Models

Probing Cross-Lingual Lexical Knowledge from Multilingual Sentence Encoders

2022-04-30 · Ivan Vulić, Goran Glavaš, Fangyu Liu, Nigel Collier 외

Pretrained multilingual language models (LMs) can be successfully transformed into multilingual sentence encoders (SEs; e.g., LaBSE, xMPNet) via additional fine-tuning or model distillation with parallel data. However, i…

Contrastive LearningCross-Lingual Entity LinkingEntity LinkingPretrained Multilingual Language Models+4

Team ÚFAL at CMCL 2022 Shared Task: Figuring out the correct recipe for predicting Eye-Tracking features using Pretrained Language Models

2022-04-11 · CMCL (ACL) 2022 5 · Sunit Bhattacharya, Rishu Kumar, Ondrej Bojar

Eye-Tracking data is a very useful source of information to study cognition and especially language comprehension in humans. In this paper, we describe our systems for the CMCL 2022 shared task on predicting eye-tracking…

Pretrained Multilingual Language Models

Improving Word Translation via Two-Stage Contrastive Learning

2022-03-15 · ACL 2022 5 · Yaoyiran Li, Fangyu Liu, Nigel Collier, Anna Korhonen 외

Word translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages. In this work, we propose a robust and effective two-stage contrastive learn…

Bilingual Lexicon InductionContrastive LearningCross-Lingual Word EmbeddingsMachine Translation+7

Out of Thin Air: Is Zero-Shot Cross-Lingual Keyword Detection Better Than Unsupervised?

2022-02-14 · LREC 2022 6 · Boshko Koloski, Senja Pollak, Blaž Škrlj, Matej Martinc

Keyword extraction is the task of retrieving words that are essential to the content of a given document. Researchers proposed various approaches to tackle this problem. At the top-most level, approaches are divided into…

Keyword ExtractionPretrained Multilingual Language Models

Are Pretrained Multilingual Models Equally Fair Across Languages?

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Pretrained multilingual language models can help bridge the digital language divide, enabling high-quality NLP models for lower-resourced languages. Studies of multilingual models have so far focused on performance, cons…

Cloze TestFairnessPretrained Multilingual Language ModelsXLM-R

Investigating Math Word Problems using Pretrained Multilingual Language Models

2022-01-16 · ACL ARR January 2022 1 · Anonymous

In this paper, we revisit math word problems~(MWPs) from the {\em cross-lingual} and {\em multilingual} perspective.We construct our MWP solvers over pretrained multilingual language models using the sequence-to-sequence…

Machine TranslationMathPretrained Multilingual Language ModelsTranslation

Multi-Source Cross-Lingual Constituency Parsing

2021-12-01 · ICON 2021 12 · Hour Kaing, Chenchen Ding, Katsuhito Sudoh, Masao Utiyama 외

Pretrained multilingual language models have become a key part of cross-lingual transfer for many natural language processing tasks, even those without bilingual information. This work further investigates the cross-ling…

Constituency ParsingCross-Lingual TransferDiversityPretrained Multilingual Language Models

Improving Word Translation via Two-Stage Contrastive Learning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Word translation or bilingual lexicon induction (BLI) is a key cross-lingual task, aiming to bridge the lexical gap between different languages. In this work, we propose a robust and effective two-stage contrastive learn…

Bilingual Lexicon InductionContrastive LearningCross-Lingual Word EmbeddingsMachine Translation+8

Small Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages

2021-11-01 · EMNLP (MRL) 2021 11 · Kelechi Ogueji, Yuxin Zhu, Jimmy Lin

Pretrained multilingual language models have been shown to work well on many languages for a variety of downstream NLP tasks. However, these models are known to require a lot of training data. This consequently leaves ou…

Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+4

Transliteration: A Simple Technique For Improving Multilingual Language Modeling

2021-09-29 · Ibraheem Muhammad Moosa, Mahmud Elahi Akhter, Ashfia Binte Habib

While impressive performance in natural language processing tasks has been achieved for many languages by transfer learning from large pretrained multilingual language models, it is limited by the unavailability of large…

Language ModelingLanguage ModellingMultiple-choiceMultiple Choice Question Answering (MCQA)+5
1–20 / 26 다음 →