Toward Better Loanword Identification in Uyghur Using Cross-lingual Word Embeddings
To enrich vocabulary of low resource settings, we proposed a novel method which identify loanwords in monolingual corpora. More specifically, we first use cross-lingual word embeddings as the core feature to generate semantically related candidates based on comparable corpora and a small bilingual lexicon; then, a log-linear model which combines several shallow features such as pronunciation similarity and hybrid language model features to predict the final results. In this paper, we use Uyghur as the receipt language and try to detect loanwords in four donor languages: Arabic, Chinese, Persian and Russian. We conduct two groups of experiments to evaluate the effectiveness of our proposed approach: loanword identification and OOV translation in four language pairs and eight translation directions (Uyghur-Arabic, Arabic-Uyghur, Uyghur-Chinese, Chinese-Uyghur, Uyghur-Persian, Persian-Uyghur, Uyghur-Russian, and Russian-Uyghur). Experimental results on loanword identification show that our method outperforms other baseline models significantly. Neural machine translation models integrating results of loanword identification experiments achieve the best results on OOV translation(with 0.5-0.9 BLEU improvements)
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual Word EmbeddingsLanguage ModelingLanguage ModellingMachine TranslationTranslationWord EmbeddingsSimilar Papers 제목 키워드 기반
Recurrent Neural Network Based Loanwords Identification in Uyghur
A Neural Network Based Model for Loanword Identification in Uyghur
Are Language Models Borrowing-Blind? A Multilingual Evaluation of Loanword Identification across 10 Languages
Throughout language history, words are borrowed from one language to another and gradually become integrated into the recipient's lexicon. Speakers can often differentiate these loanwords from native vocabulary, particul…
Automatic Speech Recognition for Uyghur through Multilingual Acoustic Modeling
Low-resource languages suffer from lower performance of Automatic Speech Recognition (ASR) system due to the lack of data. As a common approach, multilingual training has been applied to achieve more context coverage and…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionA Generalized Method for Automated Multilingual Loanword Detection
Loanwords are words incorporated from one language into another without translation. Suppose two words from distantly-related or unrelated languages sound similar and have a similar meaning. In that case, this is evidenc…
Semantic SimilaritySemantic Textual Similarity