paper-with-me

홈 › Papers

EMS: Efficient and Effective Massively Multilingual Sentence Embedding Learning

2022-05-31 · Zhuoyuan Mao, Chenhui Chu, Sadao Kurohashi

Massively multilingual sentence representation models, e.g., LASER, SBERT-distill, and LaBSE, help significantly improve cross-lingual downstream tasks. However, the use of a large amount of data or inefficient model architectures results in heavy computation to train a new model according to our preferred languages and domains. To resolve this issue, we introduce efficient and effective massively multilingual sentence embedding (EMS), using cross-lingual token-level reconstruction (XTR) and sentence-level contrastive learning as training objectives. Compared with related studies, the proposed model can be efficiently trained using significantly fewer parallel sentences and GPU computation resources. Empirical results showed that the proposed model significantly yields better or comparable results with regard to cross-lingual sentence retrieval, zero-shot cross-lingual genre classification, and sentiment classification. Ablative analyses demonstrated the efficiency and effectiveness of each component of the proposed model. We release the codes for model training and the EMS pre-trained sentence embedding model, which supports 62 languages ( https://github.com/Mao-KU/EMS ).

📄 PDF Abstract BibTeX arXiv:2205.15744

Code (1)

mao-ku/ems 공식 구현 pytorch

Tasks

Contrastive LearningGenre classificationGPURepresentation LearningRetrievalSentenceSentence EmbeddingSentence-EmbeddingSentence RetrievalSentiment AnalysisSentiment Classification

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

A General-Purpose Multilingual Document Encoder

2023-05-11 · Onur Galoğlu, Robert Litschko, Goran Glavaš

Massively multilingual pretrained transformers (MMTs) have tremendously pushed the state of the art on multilingual NLP and cross-lingual transfer of NLP models in particular. While a large body of work leveraged MMTs to…

Cross-Lingual TransferDocument ClassificationLong-range modelingMultilingual NLP+2

Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond

2018-12-26 · TACL 2019 3 · Mikel Artetxe, Holger Schwenk

We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts. Our system uses a single BiLSTM encode…

Cross-Lingual Bitext MiningCross-Lingual Document ClassificationCross-Lingual Natural Language InferenceCross-Lingual Transfer+8

News Without Borders: Domain Adaptation of Multilingual Sentence Embeddings for Cross-lingual News Recommendation

2024-06-18 · Andreea Iana, Fabian David Schmidt, Goran Glavaš, Heiko Paulheim

Rapidly growing numbers of multilingual news consumers pose an increasing challenge to news recommender systems in terms of providing customized recommendations. First, existing neural news recommenders, even when powere…

Cross-Lingual TransferDomain AdaptationMultilingual NLPNews Recommendation+8

Preparing the Vuk'uzenzele and ZA-gov-multilingual South African multilingual corpora

2023-03-07 · Richard Lastrucci, Isheanesu Dzingirai, Jenalea Rajab, Andani Madodonga 외

This paper introduces two multilingual government themed corpora in various South African languages. The corpora were collected by gathering the South African Government newspaper (Vuk'uzenzele), as well as South African…

Language ModelingLanguage ModellingMachine TranslationNMT+1

Massively Multilingual Lexical Specialization of Multilingual Transformers

2022-08-01 · Tommaso Green, Simone Paolo Ponzetto, Goran Glavaš

While pretrained language models (PLMs) primarily serve as general-purpose text encoders that can be fine-tuned for a wide variety of downstream tasks, recent work has shown that they can also be rewired to produce high-…

Bilingual Lexicon InductionRetrievalSentenceSentence Retrieval+3