An Iterative Approach for Mining Parallel Sentences in a Comparable Corpus
We describe an approach for mining parallel sentences in a collection of documents in two languages. While several approaches have been proposed for doing so, our proposal differs in several respects. First, we use a document level classifier in order to focus on potentially fruitful document pairs, an understudied approach. We show that mining less, but more parallel documents can lead to better gains in machine translation. Second, we compare different strategies for post-processing the output of a classifier trained to recognize parallel sentences. Last, we report a simple bootstrapping experiment which shows that promising sentence pairs extracted in a first stage can help to mine new sentence pairs in a second stage. We applied our approach on the English-French Wikipedia. Gains of a statistical machine translation (SMT) engine are analyzed along different test sets.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalMachine TranslationSentenceTranslationSimilar Papers 제목 키워드 기반
Harvesting comparable corpora and mining them for equivalent bilingual sentences using statistical classification and analogy- based heuristics
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for minin…
General ClassificationMachine TranslationRetrievalTranslationPEXACC: A Parallel Sentence Mining Algorithm from Comparable Corpora
Extracting parallel data from comparable corpora in order to enrich existing statistical translation models is an avenue that attracted a lot of research in recent years. There are experiments that convincingly show how …
Information RetrievalMachine TranslationSentenceTranslationParallel Sentence Mining by Constrained Decoding
We present a novel method to extract parallel sentences from two monolingual corpora, using neural machine translation. Our method relies on translating sentences in one corpus, but constraining the decoding by a prefix …
Cross-Lingual Bitext MiningMachine TranslationParallel Corpus MiningSentence+1Volctrans Parallel Corpus Filtering System for WMT 2020
In this paper, we describe our submissions to the WMT20 shared task on parallel corpus filtering and alignment for low-resource conditions. The task requires the participants to align potential parallel sentence pairs ou…
RerankingSentenceWord AlignmentA Corpus for English-Japanese Multimodal Neural Machine Translation with Comparable Sentences
Multimodal neural machine translation (NMT) has become an increasingly important area of research over the years because additional modalities, such as image data, can provide more context to textual data. Furthermore, t…
Image CaptioningMachine TranslationNMTSentence+1