paper-with-me

홈 › Papers

Collection of a Large Database of French-English SMT Output Corrections

2012-05-01 · LREC 2012 5 · Marion Potet, Emmanuelle Esperan{\c{c}}a-Rodier, Laurent Besacier, Herv{\'e} Blanchon

Corpus-based approaches to machine translation (MT) rely on the availability of parallel corpora. To produce user-acceptable translation outputs, such systems need high quality data to be efficiency trained, optimized and evaluated. However, building high quality dataset is a relatively expensive task. In this paper, we describe the data collection and analysis of a large database of 10.881 SMT translation output hypotheses manually corrected. These post-editions were collected using Amazon's Mechanical Turk, following some ethical guidelines. A complete analysis of the collected data pointed out a high quality of the corrections with more than 87 {\%} of the collected post-editions that improve hypotheses and more than 94 {\%} of the crowdsourced post-editions which are at least of professional quality. We also post-edited 1,500 gold-standard reference translations (of bilingual parallel corpora generated by professional) and noticed that 72 {\%} of these translations needed to be corrected during post-edition. We computed a proximity measure between the differents kind of translations and pointed out that reference translations are as far from the hypotheses than from the corrected hypotheses (i.e. the post-editions). In light of these last findings, we discuss the adequation of text-based generated reference translations to train setence-to-sentence based SMT systems.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

PAGnol: An Extra-Large French Generative Model

2021-10-16 · LREC 2022 6 · Julien Launay, E. L. Tommasone, Baptiste Pannier, François Boniface 외

Access to large pre-trained models of varied architectures, in many different languages, is central to the democratization of NLP. We introduce PAGnol, a collection of French GPT models. Using scaling laws, we efficientl…

model

Preserving Semantic Information from Old Dictionaries: Linking Senses of the `Altfranz\"osisches W\"orterbuch' to WordNet

2020-05-01 · LREC 2020 5 · Achim Stein

Historical dictionaries of the pre-digital period are important resources for the study of older languages. Taking the example of the {`}Altfranz{\"o}sisches W{\"o}rterbuch{'}, an Old French dictionary published from 192…

Optical Character Recognition (OCR)

The Multilingual Paraphrase Database

2014-05-01 · LREC 2014 5 · Juri Ganitkevitch, Chris Callison-Burch

We release a massive expansion of the paraphrase database (PPDB) that now includes a collection of paraphrases in 23 different languages. The resource is derived from large volumes of bilingual parallel data. Our collect…

Document SummarizationInformation RetrievalMachine TranslationMulti-Document Summarization+4

LongEval-Retrieval: French-English Dynamic Test Collection for Continuous Web Search Evaluation

2023-03-06 · Petra Galuščáková Romain Deveaud, Gabriela Gonzalez-Saez, Philippe Mulhem, Lorraine Goeuriot 외

LongEval-Retrieval is a Web document retrieval benchmark that focuses on continuous retrieval evaluation. This test collection is intended to be used to study the temporal persistence of Information Retrieval systems and…

Information RetrievalPrivacy PreservingRetrieval

Use of a Citizen Science Platform for the Creation of a Language Resource to Study Bias in Language Models for French: A Case Study

2022-06-01 · NIDCP (LREC) 2022 6 · Karën Fort, Aurélie Névéol, Yoann Dupont, Julien Bezançon

There is a growing interest in the evaluation of bias, fairness and social impact of Natural Language Processing models and tools. However, little resources are available for this task in languages other than English. Tr…

FairnessTranslation