paper-with-me

Papers

How Do Source-side Monolingual Word Embeddings Impact Neural Machine Translation?

2018-06-05 · Shuoyang Ding, Kevin Duh

Using pre-trained word embeddings as input layer is a common practice in many natural language processing (NLP) tasks, but it is largely neglected for neural machine translation (NMT). In this paper, we conducted a systematic analysis on the effect of using pre-trained source-side monolingual word embedding in NMT. We compared several strategies, such as fixing or updating the embeddings during NMT training on varying amounts of data, and we also proposed a novel strategy called dual-embedding that blends the fixing and updating strategies. Our results suggest that pre-trained embeddings can be helpful if properly incorporated into NMT, especially when parallel data is limited or additional in-domain monolingual data is readily available.

📄 PDF Abstract BibTeX arXiv:1806.01515

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTTranslationWord Embeddings

Similar Papers 제목 키워드 기반

Monolingual Embeddings for Low Resourced Neural Machine Translation

2017-12-01 · IWSLT 2017 12 · Mattia Antonino Di Gangi, Marcello Federico

Neural machine translation (NMT) is the state of the art for machine translation, and it shows the best performance when there is a considerable amount of data available. When only little data exist for a language pair, …

Machine TranslationNMTTranslationWord Embeddings

Anchor-based Bilingual Word Embeddings for Low-Resource Languages

2020-10-23 · ACL 2021 5 · Tobias Eder, Viktor Hangya, Alexander Fraser

Good quality monolingual word embeddings (MWEs) can be built for languages which have large amounts of unlabeled text. MWEs can be aligned to bilingual spaces using only a few thousand word translation pairs. For low res…

Bilingual Lexicon InductionCross-Lingual TransferTransfer LearningTranslation+3

Meemi: A Simple Method for Post-processing and Integrating Cross-lingual Word Embeddings

2019-10-16 · Yerai Doval, Jose Camacho-Collados, Luis Espinosa-Anke, Steven Schockaert

Word embeddings have become a standard resource in the toolset of any Natural Language Processing practitioner. While monolingual word embeddings encode information about words in the context of a particular language, cr…

Cross-Lingual Natural Language InferenceCross-Lingual Word EmbeddingsHypernym DiscoveryNatural Language Inference+2

Low-Resource Unsupervised NMT: Diagnosing the Problem and Providing a Linguistically Motivated Solution

2020-11-01 · EAMT 2020 11 · Lukas Edman, Antonio Toral, Gertjan van Noord

Unsupervised Machine Translation has been advancing our ability to translate without parallel data, but state-of-the-art methods assume an abundance of monolingual data. This paper investigates the scenario where monolin…

Machine TranslationNMTTranslationUnsupervised Machine Translation+1

IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages

2020-11-08 · Findings of the Association for Computational Linguistics 2020 · Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C. 외

In this paper, we introduce NLP resources for 11 major Indian languages from two major language families. These resources include: (a) large-scale sentence-level monolingual corpora, (b) pre-trained word embeddings, (c)…

Genre classificationMultiple-choiceMultiple Choice Question Answering (MCQA)named-entity-recognition+8