paper-with-me

Papers

Advancing Multilingual Pre-training: TRIP Triangular Document-level Pre-training for Multilingual Language Models

2022-12-15 · Hongyuan Lu, Haoyang Huang, Shuming Ma, Dongdong Zhang, Wai Lam, Furu Wei

Despite the success of multilingual sequence-to-sequence pre-training, most existing approaches rely on document-level monolingual corpora in many different languages, sentence-level bilingual corpora,\footnote{In this paper, we use bilingual corpora' to denote parallel corpora with bilingual translation pairs' in many different language pairs, each consisting of two sentences/documents with the same meaning written in different languages. We use trilingual corpora' to denote parallel corpora with trilingual translation pairs' in many different language combinations, each consisting of three sentences/documents.} and sometimes synthetic document-level bilingual corpora. This hampers the performance with cross-lingual document-level tasks such as document-level translation. Therefore, we propose to mine and leverage document-level trilingual parallel corpora to improve sequence-to-sequence multilingual pre-training. We present \textbf{Tri}angular Document-level \textbf{P}re-training (\textbf{TRIP}), which is the first in the field to accelerate the conventional monolingual and bilingual objectives into a trilingual objective with a novel method called Grafting. Experiments show that TRIP achieves several strong state-of-the-art (SOTA) scores on three multilingual document-level machine translation benchmarks and one cross-lingual abstractive summarization benchmark, including consistent improvements by up to 3.11 d-BLEU points and 8.9 ROUGE-L points.

📄 PDF Abstract BibTeX arXiv:2212.07752

Code (0)

등록된 구현이 없습니다.

Tasks

Abstractive Text SummarizationCross-Lingual Abstractive SummarizationDocument Level Machine TranslationMachine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

TriAdReview: Triangular Adversarial Review Architecture for Multi-Model Technical Document Generation

2026-06-13 · Zhiqiang Zhou, Junliang Dai, Xu Ling arxiv

Large language models (LLMs) are increasingly used for technical document generation, yet single-model outputs often suffer from over-engineering, security blind spots, and incomplete coverage. We propose TriAdReview, a …

Code Generation

DocHPLT: A Massively Multilingual Document-Level Translation Dataset

2025-08-18 · Dayyán O'Brien, Bhavitvya Malik, Ona de Gibert, Pinzhen Chen 외 arxiv

Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, …

Machine Translation

CLIRMatrix: A massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval

2020-11-01 · EMNLP 2020 11 · Shuo Sun, Kevin Duh

We present CLIRMatrix, a massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval extracted automatically from Wikipedia. CLIRMatrix comprises (1) BI-139, a bilingual data…

Cross-Lingual Information RetrievalInformation RetrievalRetrieval

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

2024-12-11 · Andreas Koukounas, Georgios Mastrapas, Sedigheh Eslami, Bo wang 외

Contrastive Language-Image Pretraining (CLIP) has been widely used for crossmodal information retrieval and multimodal understanding tasks. However, CLIP models are mainly optimized for crossmodal vision-language tasks a…

Contrastive LearningCross-Modal Information RetrievalInformation RetrievalRepresentation Learning+3

A Robust Self-Learning Framework for Cross-Lingual Text Classification

2019-11-01 · IJCNLP 2019 11 · Xin Dong, Gerard de Melo

Based on massive amounts of data, recent pretrained contextual representation models have made significant strides in advancing a number of different English NLP tasks. However, for other languages, relevant training dat…

ClassificationGeneral ClassificationSelf-LearningSentiment Analysis+3