paper-with-me

Papers

Collecting and Using Comparable Corpora for Statistical Machine Translation

2012-05-01 · LREC 2012 5 · Inguna Skadi{\c{n}}a, Ahmet Aker, Nikos Mastropavlos, Fangzhong Su, Dan Tufis, Mateja Verlic, Andrejs Vasi{\c{l}}jevs, Bogdan Babych, Paul Clough, Robert Gaizauskas, Nikos Glaros, Monica Lestari Paramita, M{\=a}rcis Pinnis

Lack of sufficient parallel data for many languages and domains is currently one of the major obstacles to further advancement of automated translation. The ACCURAT project is addressing this issue by researching methods how to improve machine translation systems by using comparable corpora. In this paper we present tools and techniques developed in the ACCURAT project that allow additional data needed for statistical machine translation to be extracted from comparable corpora. We present methods and tools for acquisition of comparable corpora from the Web and other sources, for evaluation of the comparability of collected corpora, for multi-level alignment of comparable corpora and for extraction of lexical and terminological data for machine translation. Finally, we present initial evaluation results on the utility of collected corpora in domain-adapted machine translation and real-life applications.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Harvesting comparable corpora and mining them for equivalent bilingual sentences using statistical classification and analogy- based heuristics

2015-11-18 · Krzysztof Wołk, Emilia Rejmund, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for minin…

General ClassificationMachine TranslationRetrievalTranslation

Creation of comparable corpora for English-Urdu, Arabic, Persian

2016-05-01 · LREC 2016 5 · Murad Abouammoh, Kashif Shah, Ahmet Aker

Statistical Machine Translation (SMT) relies on the availability of rich parallel corpora. However, in the case of under-resourced languages or some specific domains, parallel corpora are not readily available. This lead…

ArticlesMachine TranslationTranslation

PEXACC: A Parallel Sentence Mining Algorithm from Comparable Corpora

2012-05-01 · LREC 2012 5 · Radu Ion

Extracting parallel data from comparable corpora in order to enrich existing statistical translation models is an avenue that attracted a lot of research in recent years. There are experiments that convincingly show how …

Information RetrievalMachine TranslationSentenceTranslation

Building Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs

2015-09-29 · Krzysztof Wołk, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such…

ArticlesMachine TranslationRetrievalSentence+1

Polish - English Speech Statistical Machine Translation Systems for the IWSLT 2014

2015-09-29 · Krzysztof Wołk, Krzysztof Marasek

This research explores effects of various training settings between Polish and English Statistical Machine Translation systems for spoken language. Various elements of the TED parallel text corpora for the IWSLT 2014 eva…

LEMMAMachine TranslationTranslation