paper-with-me

Papers

A Richly Annotated, Multilingual Parallel Corpus for Hybrid Machine Translation

2012-05-01 · LREC 2012 5 · Eleftherios Avramidis, Marta R. Costa-juss{\`a}, Christian Federmann, Josef van Genabith, Maite Melero, Pavel Pecina

In recent years, machine translation (MT) research has focused on investigating how hybrid machine translation as well as system combination approaches can be designed so that the resulting hybrid translations show an improvement over the individual “component” translations. As a first step towards achieving this objective we have developed a parallel corpus with source text and the corresponding translation output from a number of machine translation engines, annotated with metadata information, capturing aspects of the translation process performed by the different MT systems. This corpus aims to serve as a basic resource for further research on whether hybrid machine translation algorithms and system combination techniques can benefit from additional (linguistically motivated, decoding, and runtime) information provided by the different systems involved. In this paper, we describe the annotated corpus we have created. We provide an overview on the component MT systems and the XLIFF-based annotation format we have developed. We also report on first experiments with the ML4HMT corpus data.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration

2020-05-01 · LREC 2020 5 · Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller 외

We present findings from the creation of a massively parallel corpus in over 1600 languages, the Johns Hopkins University Bible Corpus (JHUBC). The corpus consists of over 4000 unique translations of the Christian Bible …

HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models

2025-11-02 · Stephan Oepen, Nikolay Arefev, Mikko Aulamo, Marta Bañón 외 arxiv

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingu…

Language IdentificationMachine Translation

Building The Sense-Tagged Multilingual Parallel Corpus

2014-05-01 · LREC 2014 5 · Shan Wang, Francis Bond

Sense-annotated parallel corpora play a crucial role in natural language processing. This paper introduces our progress in creating such a corpus for Asian languages using English as a pivot, which is the first such corp…

The making of the Litkey Corpus, a richly annotated longitudinal corpus of German texts written by primary school children

2019-08-01 · WS 2019 8 · Ronja Laarmann-Quante, Stefanie Dipper, Eva Belke

To date, corpus and computational linguistic work on written language acquisition has mostly dealt with second language learners who have usually already mastered orthography acquisition in their first language. In this …

Language AcquisitionPOS

Multilingual Sense Intersection in a Parallel Corpus with Diverse Language Families

2016-01-01 · GWC 2016 1 · Giulia Bonansinga, Francis Bond

Supervised methods for Word Sense Disambiguation (WSD) benefit from high-quality sense-annotated resources, which are lacking for many languages less common than English. There are, however, several multilingual parallel…

Word Sense Disambiguation