Cross-Lingual Bitext Mining
4개 벤치마크 · 논문 6편 · 이 태스크의 논문 보기 →
Benchmarks
BUCC French-to-English
BUCC German-to-English
BUCC Chinese-to-English
BUCC Russian-to-English
Most implemented
Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings
Improving Neural Machine Translation Models with Monolingual Data
Majority Voting with Bidirectional Pre-translation For Bitext Retrieval
Parallel Sentence Mining by Constrained Decoding
Papers
Low-Resource Machine Translation Training Curriculum Fit for Low-Resource Languages
We conduct an empirical study of neural machine translation (NMT) for truly low-resource languages, and propose a training curriculum fit for cases when both parallel training data and compute resource are lacking, refle…
Cross-Lingual Bitext MiningLanguage ModellingLow-Resource Neural Machine TranslationMachine Translation+2Majority Voting with Bidirectional Pre-translation For Bitext Retrieval
Obtaining high-quality parallel corpora is of paramount importance for training NMT systems. However, as many language pairs lack adequate gold-standard training data, a popular approach has been to mine so-called "pseud…
Cross-Lingual Bitext MiningNMTRetrievalTranslationParallel Sentence Mining by Constrained Decoding
We present a novel method to extract parallel sentences from two monolingual corpora, using neural machine translation. Our method relies on translating sentences in one corpus, but constraining the decoding by a prefix …
Cross-Lingual Bitext MiningMachine TranslationParallel Corpus MiningSentence+1Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts. Our system uses a single BiLSTM encode…
Cross-Lingual Bitext MiningCross-Lingual Document ClassificationCross-Lingual Natural Language InferenceCross-Lingual Transfer+8Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings
Machine translation is highly sensitive to the size and quality of the training data, which has led to an increasing interest in collecting and filtering large parallel corpora. In this paper, we propose a new method for…
Cross-Lingual Bitext MiningMachine TranslationParallel Corpus MiningRetrieval+3Improving Neural Machine Translation Models with Monolingual Data
Neural Machine Translation (NMT) has obtained state-of-the art performance for several language pairs, while only using parallel data for training. Target-side monolingual data plays an important role in boosting fluency…
Cross-Lingual Bitext MiningDecoderLanguage ModelingLanguage Modelling+3