paper-with-me

Papers

The United Nations Parallel Corpus v1.0

2016-05-01 · LREC 2016 5 · Micha{\l} Ziemski, Marcin Junczys-Dowmunt, Bruno Pouliquen

This paper describes the creation process and statistics of the official United Nations Parallel Corpus, the first parallel corpus composed from United Nations documents published by the original data creator. The parallel corpus presented consists of manually translated UN documents from the last 25 years (1990 to 2014) for the six official UN languages, Arabic, Chinese, English, French, Russian, and Spanish. The corpus is freely available for download under a liberal license. Apart from the pairwise aligned documents, a fully aligned subcorpus for the six official UN languages is distributed. We provide baseline BLEU scores of our Moses-based SMT systems trained with the full data of language pairs involving English and for all possible translation directions of the six-way subcorpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

UPRPRC: Unified Pipeline for Reproducing Parallel Resources -- Corpus from the United Nations

2025-09-19 · Qiuyang Lu, Fangjian Shen, Zhengkai Tang, Qiang Liu 외 arxiv

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, diffic…

Machine Translation

MultiUN v2: UN Documents with Multilingual Alignments

2012-05-01 · LREC 2012 5 · Yu Chen, Andreas Eisele

MultiUN is a multilingual parallel corpus extracted from the official documents of the United Nations. It is available in the six official languages of the UN and a small portion of it is also available in German. This p…

Information RetrievalMachine TranslationSentenceTranslation

Effective Parallel Corpus Mining using Bilingual Sentence Embeddings

2018-07-31 · WS 2018 10 · Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge 외

This paper presents an effective approach for parallel corpus mining using bilingual sentence embeddings. Our embedding models are trained to produce similar representations exclusively for bilingual sentence pairs that …

Machine TranslationNMTParallel Corpus MiningSemantic Similarity+4

Monolingual and Parallel Corpora for Kangri Low Resource Language

2021-03-22 · Shweta Chauhan, Shefali Saxena, Philemon Daniel

In this paper we present the dataset of Himachali low resource endangered language, Kangri (ISO 639-3xnr) listed in the United Nations Educational, Scientific and Cultural Organization (UNESCO). The compilation of kangri…

Machine TranslationNMTTranslationWord Embeddings

Innovations in Parallel Corpus Search Tools

2014-05-01 · LREC 2014 5 · Martin Volk, Johannes Gra{\"e}n, Elena Callegaro

Recent years have seen an increased interest in and availability of parallel corpora. Large corpora from international organizations (e.g. European Union, United Nations, European Patent Office), or from multilingual Int…

Machine TranslationSentenceTranslationWord Alignment+1