paper-with-me

Papers

OpusTools and Parallel Corpus Diagnostics

2020-05-01 · LREC 2020 5 · Mikko Aulamo, Umut Sulubacak, Sami Virpioja, J{\"o}rg Tiedemann

This paper introduces OpusTools, a package for downloading and processing parallel corpora included in the OPUS corpus collection. The package implements tools for accessing compressed data in their archived release format and make it possible to easily convert between common formats. OpusTools also includes tools for language identification and data filtering as well as tools for importing data from various sources into the OPUS format. We show the use of these tools in parallel corpus creation and data diagnostics. The latter is especially useful for the identification of potential problems and errors in the extensive data set. Using these tools, we can now monitor the validity of data sets and improve the overall quality and consitency of the data collection.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identification

Similar Papers 제목 키워드 기반

The IIT Bombay English-Hindi Parallel Corpus

2017-10-08 · LREC 2018 5 · Anoop Kunchukuttan, Pratik Mehta, Pushpak Bhattacharyya

We present the IIT Bombay English-Hindi Parallel Corpus. The corpus is a compilation of parallel corpora previously available in the public domain as well as new parallel corpora we collected. The corpus contains 1.49 mi…

Machine TranslationNMTTranslation

Alibaba Submission to the WMT18 Parallel Corpus Filtering Task

2018-10-01 · WS 2018 10 · Jun Lu, Xiaoyu Lv, Yangbin Shi, Boxing Chen

This paper describes the Alibaba Machine Translation Group submissions to the WMT 2018 Shared Task on Parallel Corpus Filtering. While evaluating the quality of the parallel corpus, the three characteristics of the corpu…

DiversityMachine TranslationSentenceTranslation+1

Automatic Parallel Corpus Creation for Hindi-English News Translation Task

2019-01-24 · Aditya Kumar Pathak, Priyankit Acharya, Dilpreet Kaur, Rakesh Chandra Balabantaray

The parallel corpus for multilingual NLP tasks, deep learning applications like Statistical Machine Translation Systems is very important. The parallel corpus of Hindi-English language pair available for news translation…

Machine TranslationMultilingual NLPTranslation

Machine Translation Model based on Non-parallel Corpus and Semi-supervised Transductive Learning

2014-05-22 · Lijiang Chen

Although the parallel corpus has an irreplaceable role in machine translation, its scale and coverage is still beyond the actual needs. Non-parallel corpus resources on the web have an inestimable potential value in mach…

Machine TranslationTransductive LearningTranslation

JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus

2022-02-25 · LREC 2022 6 · Makoto Morishita, Katsuki Chousa, Jun Suzuki, Masaaki Nagata

Most current machine translation models are mainly trained with parallel corpora, and their translation accuracy largely depends on the quality and quantity of the corpora. Although there are billions of parallel sentenc…

Machine TranslationSentenceTranslation