paper-with-me

홈 › Papers

The ILSP/ARC submission to the WMT 2018 Parallel Corpus Filtering Shared Task

2018-10-01 · WS 2018 10 · Vassilis Papavassiliou, Sokratis Sofianopoulos, Prokopis Prokopidis, Stelios Piperidis

This paper describes the submission of the Institute for Language and Speech Processing/Athena Research and Innovation Center (ILSP/ARC) for the WMT 2018 Parallel Corpus Filtering shared task. We explore several properties of sentences and sentence pairs that our system explored in the context of the task with the purpose of clustering sentence pairs according to their appropriateness in training MT systems. We also discuss alternative methods for ranking the sentence pairs of the most appropriate clusters with the aim of generating the two datasets (of 10 and 100 million words as required in the task) that were evaluated. By summarizing the results of several experiments that were carried out by the organizers during the evaluation phase, our submission achieved an average BLEU score of 26.41, even though it does not make use of any language-specific resources like bilingual lexica, monolingual corpora, or MT output, while the average score of the best participant system was 27.91.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ARCClusteringLanguage ModelingLanguage ModellingMachine TranslationOutlier DetectionSentence

Similar Papers 제목 키워드 기반

Alibaba Submission to the WMT18 Parallel Corpus Filtering Task

2018-10-01 · WS 2018 10 · Jun Lu, Xiaoyu Lv, Yangbin Shi, Boxing Chen

This paper describes the Alibaba Machine Translation Group submissions to the WMT 2018 Shared Task on Parallel Corpus Filtering. While evaluating the quality of the parallel corpus, the three characteristics of the corpu…

DiversityMachine TranslationSentenceTranslation+1

Prompsit's submission to WMT 2018 Parallel Corpus Filtering shared task

2018-10-01 · WS 2018 10 · V{\'\i}ctor M. S{\'a}nchez-Cartagena, Marta Ba{\~n}{\'o}n, Sergio Ortiz-Rojas, Gema Ram{\'\i}rez

This paper describes Prompsit Language Engineering{'}s submissions to the WMT 2018 parallel corpus filtering shared task. Our four submissions were based on an automatic classifier for identifying pairs of sentences that…

Active LearningLanguage ModelingLanguage ModellingMachine Translation

Bicleaner at WMT 2020: Universitat d’Alacant-Prompsit’s submission to the parallel corpus filtering shared task

2020-11-01 · WMT (EMNLP) 2020 11 · Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Jaume Zaragoza-Bernabeu, Felipe Sánchez-Martínez

This paper describes the joint submission of Universitat d’Alacant and Prompsit Language Engineering to the WMT 2020 shared task on parallel corpus filtering. Our submission, based on the free/open-source tool Bicleaner,…

MAJE Submission to the WMT2018 Shared Task on Parallel Corpus Filtering

2018-10-01 · WS 2018 10 · Marina Fomicheva, Jes{\'u}s Gonz{\'a}lez-Rubio

This paper describes the participation of Webinterpret in the shared task on parallel corpus filtering at the Third Conference on Machine Translation (WMT 2018). The paper describes the main characteristics of our approa…

Machine TranslationTranslation

Webinterpret Submission to the WMT2019 Shared Task on Parallel Corpus Filtering

2019-08-01 · WS 2019 8 · Jes{\'u}s Gonz{\'a}lez-Rubio

This document describes the participation of Webinterpret in the shared task on parallel corpus filtering at the Fourth Conference on Machine Translation (WMT 2019). Here, we describe the main characteristics of our appr…

Machine TranslationTranslation