paper-with-me

Cross-Lingual Bitext Mining

4개 벤치마크 · 논문 6편 · 이 태스크의 논문 보기 →

Benchmarks

Most implemented

Papers

Low-Resource Machine Translation Training Curriculum Fit for Low-Resource Languages

2021-03-24 · Garry Kuwanto, Afra Feyza Akyürek, Isidora Chara Tourni, Siyang Li 외

We conduct an empirical study of neural machine translation (NMT) for truly low-resource languages, and propose a training curriculum fit for cases when both parallel training data and compute resource are lacking, refle…

Cross-Lingual Bitext MiningLanguage ModellingLow-Resource Neural Machine TranslationMachine Translation+2

Majority Voting with Bidirectional Pre-translation For Bitext Retrieval

2021-03-10 · RANLP (BUCC) 2021 9 · Alex Jones, Derry Tanti Wijaya

Obtaining high-quality parallel corpora is of paramount importance for training NMT systems. However, as many language pairs lack adequate gold-standard training data, a popular approach has been to mine so-called "pseud…

Cross-Lingual Bitext MiningNMTRetrievalTranslation

Parallel Sentence Mining by Constrained Decoding

2020-07-01 · ACL 2020 6 · Pin-zhen Chen, Nikolay Bogoychev, Kenneth Heafield, Faheem Kirefu

We present a novel method to extract parallel sentences from two monolingual corpora, using neural machine translation. Our method relies on translating sentences in one corpus, but constraining the decoding by a prefix …

Cross-Lingual Bitext MiningMachine TranslationParallel Corpus MiningSentence+1

Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond

2018-12-26 · TACL 2019 3 · Mikel Artetxe, Holger Schwenk

We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts. Our system uses a single BiLSTM encode…

Cross-Lingual Bitext MiningCross-Lingual Document ClassificationCross-Lingual Natural Language InferenceCross-Lingual Transfer+8

Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings

2018-11-03 · ACL 2019 7 · Mikel Artetxe, Holger Schwenk

Machine translation is highly sensitive to the size and quality of the training data, which has led to an increasing interest in collecting and filtering large parallel corpora. In this paper, we propose a new method for…

Cross-Lingual Bitext MiningMachine TranslationParallel Corpus MiningRetrieval+3

Improving Neural Machine Translation Models with Monolingual Data

2015-11-20 · ACL 2016 8 · Rico Sennrich, Barry Haddow, Alexandra Birch

Neural Machine Translation (NMT) has obtained state-of-the art performance for several language pairs, while only using parallel data for training. Target-side monolingual data plays an important role in boosting fluency…

Cross-Lingual Bitext MiningDecoderLanguage ModelingLanguage Modelling+3