paper-with-me

Papers

ParaPat: The Multi-Million Sentences Parallel Corpus of Patents Abstracts

2020-05-01 · LREC 2020 5 · Felipe Soares, Mark Stevenson, Diego Bartolome, Anna Zaretskaya

The Google Patents is one of the main important sources of patents information. A striking characteristic is that many of its abstracts are presented in more than one language, thus making it a potential source of parallel corpora. This article presents the development of a parallel corpus from the open access Google Patents dataset in 74 language pairs, comprising more than 68 million sentences and 800 million tokens. Sentences were automatically aligned using the Hunalign algorithm for the largest 22 language pairs, while the others were abstract (i.e. paragraph) aligned. We demonstrate the capabilities of our corpus by training Neural Machine Translation (NMT) models for the main 9 language pairs, with a total of 18 models. Our parallel corpus is freely available in TSV format and with a SQLite database, with complementary information regarding patent metadata.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus

2025-08-22 · Masaaki Nagata, Katsuki Chousa, Norihito Yasuda arxiv

We constructed JaParaPat (Japanese-English Parallel Patent Application Corpus), a bilingual corpus of more than 300 million Japanese-English sentence pairs from patent applications published in Japan and the United State…

HindEnCorp - Hindi-English and Hindi-only Corpus for Machine Translation

2014-05-01 · LREC 2014 5 · Ond{\v{r}}ej Bojar, Vojt{\v{e}}ch Diatka, Pavel Rychl{\'y}, Pavel Stra{\v{n}}{\'a}k 외

We present HindEnCorp, a parallel corpus of Hindi and English, and HindMonoCorp, a monolingual corpus of Hindi in their release version 0.5. Both corpora were collected from web sources and preprocessed primarily for the…

Machine TranslationTranslation

Parallel Corpus Filtering Based on Fuzzy String Matching

2019-08-01 · WS 2019 8 · Sukanta Sen, Asif Ekbal, Pushpak Bhattacharyya

In this paper, we describe the IIT Patna{'}s submission to WMT 2019 shared task on parallel corpus filtering. This shared task asks the participants to develop methods for scoring each parallel sentence from a given nois…

NMTSentence

ASPEC: Asian Scientific Paper Excerpt Corpus

2016-05-01 · LREC 2016 5 · Toshiaki Nakazawa, Manabu Yaguchi, Kiyotaka Uchimoto, Masao Utiyama 외

In this paper, we describe the details of the ASPEC (Asian Scientific Paper Excerpt Corpus), which is the first large-size parallel corpus of scientific paper domain. ASPEC was constructed in the Japanese-Chinese machine…

Machine TranslationTranslation

UM-Corpus: A Large English-Chinese Parallel Corpus for Statistical Machine Translation

2014-05-01 · LREC 2014 5 · Liang Tian, Derek F. Wong, Lidia S. Chao, Paulo Quaresma 외

Parallel corpus is a valuable resource for cross-language information retrieval and data-driven natural language processing systems, especially for Statistical Machine Translation (SMT). However, most existing parallel c…

Boundary DetectionDomain AdaptationInformation RetrievalMachine Translation+2