paper-with-me

홈 › Papers

Improving Word Alignment of Rare Words with Word Embeddings

2016-12-01 · COLING 2016 12 · Masoud Jalili Sabet, Heshaam Faili, Gholamreza Haffari

We address the problem of inducing word alignment for language pairs by developing an unsupervised model with the capability of getting applied to other generative alignment models. We approach the task by: i)proposing a new alignment model based on the IBM alignment model 1 that uses vector representation of words, and ii)examining the use of similar source words to overcome the problem of rare source words and improving the alignments. We apply our method to English-French corpora and run the experiments with different sizes of sentence pairs. Our results show competitive performance against the baseline and in some cases improve the results up to 6.9{\%} in terms of precision.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceWord AlignmentWord Embeddings

Similar Papers 제목 키워드 기반

Learning to Compute Word Embeddings On the Fly

2017-06-01 · ICLR 2018 1 · Dzmitry Bahdanau, Tom Bosc, Stanisław Jastrzębski, Edward Grefenstette 외

Words in natural language follow a Zipfian distribution whereby some words are frequent but most are rare. Learning representations for words in the "long tail" of this distribution requires enormous amounts of data. Rep…

Language ModelingLanguage ModellingNatural Language InferenceQuestion Answering+2

Evaluating bilingual word embeddings on the long tail

2018-06-01 · NAACL 2018 6 · Fabienne Braune, Viktor Hangya, Tobias Eder, Alex Fraser 외

Bilingual word embeddings are useful for bilingual lexicon induction, the task of mining translations of given words. Many studies have shown that bilingual word embeddings perform well for bilingual lexicon induction bu…

Bilingual Lexicon InductionMachine TranslationWord Embeddings

Sub-Word Similarity based Search for Embeddings: Inducing Rare-Word Embeddings for Word Similarity Tasks and Language Modelling

2016-12-01 · COLING 2016 12 · Mittul Singh, Clayton Greenberg, Youssef Oualil, Dietrich Klakow

Training good word embeddings requires large amounts of data. Out-of-vocabulary words will still be encountered at test-time, leaving these words without embeddings. To overcome this lack of embeddings for rare words, ex…

Language ModelingLanguage ModellingMorphological AnalysisWord Embeddings+1

Learning Embeddings for Rare Words Leveraging Internet Search Engine and Spatial Location Relationships

2021-08-01 · Joint Conference on Lexical and Computational Semantics 2021 · Xiaotao Li, Shujuan You, Yawen Niu, Wai Chen

Word embedding techniques depend heavily on the frequencies of words in the corpus, and are negatively impacted by failures in providing reliable representations for low-frequency words or unseen words during training. T…

Word Embeddings

WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words

2016-05-01 · LREC 2016 5 · Luisa Bentivogli, Mauro Cettolo, M. Amin Farajian, Marcello Federico

This paper presents WAGS (Word Alignment Gold Standard), a novel benchmark which allows extensive evaluation of WA tools on out-of-vocabulary (OOV) and rare words. WAGS is a subset of the Common Test section of the Europ…

SentenceWord Alignment