paper-with-me

홈 › Papers

UM-Corpus: A Large English-Chinese Parallel Corpus for Statistical Machine Translation

2014-05-01 · LREC 2014 5 · Liang Tian, Derek F. Wong, Lidia S. Chao, Paulo Quaresma, Francisco Oliveira, Yi Lu, Shuo Li, Yiming Wang, Long-Yue Wang

Parallel corpus is a valuable resource for cross-language information retrieval and data-driven natural language processing systems, especially for Statistical Machine Translation (SMT). However, most existing parallel corpora to Chinese are subject to in-house use, while others are domain specific and limited in size. To a certain degree, this limits the SMT research. This paper describes the acquisition of a large scale and high quality parallel corpora for English and Chinese. The corpora constructed in this paper contain about 15 million English-Chinese (E-C) parallel sentences, and more than 2 million training data and 5,000 testing sentences are made publicly available. Different from previous work, the corpus is designed to embrace eight different domains. Some of them are further categorized into different topics. The corpus will be released to the research community, which is available at the NLP2CT website.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Boundary DetectionDomain AdaptationInformation RetrievalMachine TranslationRetrievalTranslation

Similar Papers 제목 키워드 기반

NEJM-enzh: A Parallel Corpus for English-Chinese Translation in the Biomedical Domain

2020-05-18 · Boxiang Liu, Liang Huang

Machine translation requires large amounts of parallel text. While such datasets are abundant in domains such as newswire, they are less accessible in the biomedical domain. Chinese and English are two of the most widely…

Machine TranslationSentenceTranslation

ASPEC: Asian Scientific Paper Excerpt Corpus

2016-05-01 · LREC 2016 5 · Toshiaki Nakazawa, Manabu Yaguchi, Kiyotaka Uchimoto, Masao Utiyama 외

In this paper, we describe the details of the ASPEC (Asian Scientific Paper Excerpt Corpus), which is the first large-size parallel corpus of scientific paper domain. ASPEC was constructed in the Japanese-Chinese machine…

Machine TranslationTranslation

CUTE: A Multilingual Dataset for Enhancing Cross-Lingual Knowledge Transfer in Low-Resource Languages

2025-09-21 · Wenhao Zhuang, Yuan Sun arxiv

Large Language Models (LLMs) demonstrate exceptional zero-shot capabilities in various NLP tasks, significantly enhancing user experience and efficiency. However, this advantage is primarily limited to resource-rich lang…

Cross-Lingual TransferMachine Translation

The United Nations Parallel Corpus v1.0

2016-05-01 · LREC 2016 5 · Micha{\l} Ziemski, Marcin Junczys-Dowmunt, Bruno Pouliquen

This paper describes the creation process and statistics of the official United Nations Parallel Corpus, the first parallel corpus composed from United Nations documents published by the original data creator. The parall…

Translation

Improving Phrase Translation Based on Sentence Alignment of Chinese-English Parallel Corpus

2020-09-01 · ROCLING 2020 9 · Yi-Jyun Chen, Ching-Yu Helen Yang, Jason S. Chang
SentenceTranslation