Advancing Multilingual Pre-training: TRIP Triangular Document-level Pre-training for Multilingual Language Models
Despite the success of multilingual sequence-to-sequence pre-training, most existing approaches rely on document-level monolingual corpora in many different languages, sentence-level bilingual corpora,\footnote{In this paper, we use bilingual corpora' to denote parallel corpora with bilingual translation pairs' in many different language pairs, each consisting of two sentences/documents with the same meaning written in different languages. We use trilingual corpora' to denote parallel corpora with trilingual translation pairs' in many different language combinations, each consisting of three sentences/documents.} and sometimes synthetic document-level bilingual corpora. This hampers the performance with cross-lingual document-level tasks such as document-level translation. Therefore, we propose to mine and leverage document-level trilingual parallel corpora to improve sequence-to-sequence multilingual pre-training. We present \textbf{Tri}angular Document-level \textbf{P}re-training (\textbf{TRIP}), which is the first in the field to accelerate the conventional monolingual and bilingual objectives into a trilingual objective with a novel method called Grafting. Experiments show that TRIP achieves several strong state-of-the-art (SOTA) scores on three multilingual document-level machine translation benchmarks and one cross-lingual abstractive summarization benchmark, including consistent improvements by up to 3.11 d-BLEU points and 8.9 ROUGE-L points.
Code (0)
등록된 구현이 없습니다.
Tasks
Abstractive Text SummarizationCross-Lingual Abstractive SummarizationDocument Level Machine TranslationMachine TranslationSentenceTranslationSimilar Papers 제목 키워드 기반
TriAdReview: Triangular Adversarial Review Architecture for Multi-Model Technical Document Generation
Large language models (LLMs) are increasingly used for technical document generation, yet single-model outputs often suffer from over-engineering, security blind spots, and incomplete coverage. We propose TriAdReview, a …
Code GenerationDocHPLT: A Massively Multilingual Document-Level Translation Dataset
Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, …
Machine TranslationCLIRMatrix: A massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval
We present CLIRMatrix, a massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval extracted automatically from Wikipedia. CLIRMatrix comprises (1) BI-139, a bilingual data…
Cross-Lingual Information RetrievalInformation RetrievalRetrievaljina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
Contrastive Language-Image Pretraining (CLIP) has been widely used for crossmodal information retrieval and multimodal understanding tasks. However, CLIP models are mainly optimized for crossmodal vision-language tasks a…
Contrastive LearningCross-Modal Information RetrievalInformation RetrievalRepresentation Learning+3A Robust Self-Learning Framework for Cross-Lingual Text Classification
Based on massive amounts of data, recent pretrained contextual representation models have made significant strides in advancing a number of different English NLP tasks. However, for other languages, relevant training dat…
ClassificationGeneral ClassificationSelf-LearningSentiment Analysis+3