paper-with-me

Papers

PEXACC: A Parallel Sentence Mining Algorithm from Comparable Corpora

2012-05-01 · LREC 2012 5 · Radu Ion

Extracting parallel data from comparable corpora in order to enrich existing statistical translation models is an avenue that attracted a lot of research in recent years. There are experiments that convincingly show how parallel data extracted from comparable corpora is able to improve statistical machine translation. Yet, the existing body of research on parallel sentence mining from comparable corpora does not take into account the degree of comparability of the corpus being processed or the computation time it takes to extract parallel sentences from a corpus of a given size. We will show that the performance of a parallel sentence extractor crucially depends on the degree of comparability such that it is more difficult to process a weakly comparable corpus than a strongly comparable corpus. In this paper we describe PEXACC, a distributed (running on multiple CPUs), trainable parallel sentence/phrase extractor from comparable corpora. PEXACC is freely available for download with the ACCURAT Toolkit, a collection of MT-related tools developed in the ACCURAT project.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalMachine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

Building Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs

2015-09-29 · Krzysztof Wołk, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such…

ArticlesMachine TranslationRetrievalSentence+1

Harvesting comparable corpora and mining them for equivalent bilingual sentences using statistical classification and analogy- based heuristics

2015-11-18 · Krzysztof Wołk, Emilia Rejmund, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for minin…

General ClassificationMachine TranslationRetrievalTranslation

Unsupervised Parallel Sentence Extraction from Comparable Corpora

2018-10-01 · IWSLT (EMNLP) 2018 10 · Viktor Hangya, Fabienne Braune, Yuliya Kalasouskaya, Alexander Fraser

Mining parallel sentences from comparable corpora is of great interest for many downstream tasks. In the BUCC 2017 shared task, systems performed well by training on gold standard parallel sentences. However, we often wa…

SentenceWord Embeddings

Unsupervised Parallel Sentence Extraction with Parallel Segment Detection Helps Machine Translation

2019-07-01 · ACL 2019 7 · Viktor Hangya, Alex Fraser, er

Mining parallel sentences from comparable corpora is important. Most previous work relies on supervised systems, which are trained on parallel data, thus their applicability is problematic in low-resource scenarios. Rece…

Machine TranslationSentenceTranslationWord Embeddings

A Japanese-Chinese Parallel Corpus Using Crowdsourcing for Web Mining

2024-05-15 · Masaaki Nagata, Makoto Morishita, Katsuki Chousa, Norihito Yasuda

Using crowdsourcing, we collected more than 10,000 URL pairs (parallel top page pairs) of bilingual websites that contain parallel documents and created a Japanese-Chinese parallel corpus of 4.6M sentence pairs from thes…

SentenceTranslationWord Translation