paper-with-me

홈 › Papers

Learning Bilingual Sentence Embeddings via Autoencoding and Computing Similarities with a Multilayer Perceptron

2019-06-05 · WS 2019 8 · Yunsu Kim, Hendrik Rosendahl, Nick Rossenbach, Jan Rosendahl, Shahram Khadivi, Hermann Ney

We propose a novel model architecture and training algorithm to learn bilingual sentence embeddings from a combination of parallel and monolingual data. Our method connects autoencoding and neural machine translation to force the source and target sentence embeddings to share the same space without the help of a pivot language or an additional transformation. We train a multilayer perceptron on top of the sentence embeddings to extract good bilingual sentence pairs from nonparallel or noisy parallel data. Our approach shows promising performance on sentence alignment recovery and the WMT 2018 parallel corpus filtering tasks with only a single model.

📄 PDF Abstract BibTeX arXiv:1906.01942

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceSentence EmbeddingsTranslation

Similar Papers 제목 키워드 기반

An Unsupervised System for Parallel Corpus Filtering

2018-10-01 · WS 2018 10 · Viktor Hangya, Alex Fraser, er

In this paper we describe LMU Munich{'}s submission for the \textit{WMT 2018 Parallel Corpus Filtering} shared task which addresses the problem of cleaning noisy parallel corpora. The task of mining and cleaning parallel…

Domain AdaptationLanguage ModelingLanguage ModellingMachine Translation+5

Unsupervised Parallel Sentence Extraction with Parallel Segment Detection Helps Machine Translation

2019-07-01 · ACL 2019 7 · Viktor Hangya, Alex Fraser, er

Mining parallel sentences from comparable corpora is important. Most previous work relies on supervised systems, which are trained on parallel data, thus their applicability is problematic in low-resource scenarios. Rece…

Machine TranslationSentenceTranslationWord Embeddings

Effective Parallel Corpus Mining using Bilingual Sentence Embeddings

2018-07-31 · WS 2018 10 · Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge 외

This paper presents an effective approach for parallel corpus mining using bilingual sentence embeddings. Our embedding models are trained to produce similar representations exclusively for bilingual sentence pairs that …

Machine TranslationNMTParallel Corpus MiningSemantic Similarity+4

En-Ar Bilingual Word Embeddings without Word Alignment: Factors Effects

2019-08-01 · WS 2019 8 · Taghreed Alqaisi, Simon O{'}Keefe

This paper introduces the first attempt to investigate morphological segmentation on En-Ar bilingual word embeddings using bilingual word embeddings model without word alignment (BilBOWA). We investigate the effect of se…

SegmentationSentenceWord AlignmentWord Embeddings

Learning Bilingual Word Embeddings Using Lexical Definitions

2019-06-21 · WS 2019 8 · Weijia Shi, Muhao Chen, Yingtao Tian, Kai-Wei Chang

Bilingual word embeddings, which representlexicons of different languages in a shared em-bedding space, are essential for supporting se-mantic and knowledge transfers in a variety ofcross-lingual NLP tasks. Existing appr…

TranslationWord AlignmentWord Embeddings