paper-with-me

홈 › Papers

CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs

2019-11-10 · EMNLP 2020 11 · Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzman, Philipp Koehn

Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs. We mine sixty-eight snapshots of the Common Crawl corpus and identify web document pairs that are translations of each other. We release a new web dataset consisting of over 392 million URL pairs from Common Crawl covering documents in 8144 language pairs of which 137 pairs include English. In addition to curating this massive dataset, we introduce baseline methods that leverage cross-lingual representations to identify aligned documents based on their textual content. Finally, we demonstrate the value of this parallel documents dataset through a downstream task of mining parallel sentences and measuring the quality of machine translations from models trained on this mined data. Our objective in releasing this dataset is to foster new research in cross-lingual NLP across a variety of low, medium, and high-resource languages.

📄 PDF Abstract BibTeX arXiv:1911.06154

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CLIRMatrix: A massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval

2020-11-01 · EMNLP 2020 11 · Shuo Sun, Kevin Duh

We present CLIRMatrix, a massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval extracted automatically from Wikipedia. CLIRMatrix comprises (1) BI-139, a bilingual data…

Cross-Lingual Information RetrievalInformation RetrievalRetrieval

A General-Purpose Multilingual Document Encoder

2023-05-11 · Onur Galoğlu, Robert Litschko, Goran Glavaš

Massively multilingual pretrained transformers (MMTs) have tremendously pushed the state of the art on multilingual NLP and cross-lingual transfer of NLP models in particular. While a large body of work leveraged MMTs to…

Cross-Lingual TransferDocument ClassificationLong-range modelingMultilingual NLP+2

SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information Retrieval

2026-06-17 · Youngjoon Jang, Seongtae Hong, Hyeonseok Moon, Heuiseok Lim arxiv

With the rapid expansion of massive multilingual corpora, Multilingual Information Retrieval (MLIR) has emerged as a critical technology for global information access. MLIR enables users to retrieve semantically relevant…

Information Retrieval

A New Massive Multilingual Dataset for High-Performance Language Technologies

2024-03-20 · Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón 외

We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from CommonCrawl and previously unused web cr…

Language ModelingLanguage ModellingMachine TranslationManagement+1

An Expanded Massive Multilingual Dataset for High-Performance Language Technologies

2025-03-13 · Laurie Burchell, Ona de Gibert, Nikolay Arefyev, Mikko Aulamo 외

Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In this work, we present HPLT v2, a collectio…

Machine TranslationSentence