Query Translation for Cross-Language Information Retrieval using Multilingual Word Clusters
In Cross-Language Information Retrieval, finding the appropriate translation of the source language query has always been a difficult problem to solve. We propose a technique towards solving this problem with the help of multilingual word clusters obtained from multilingual word embeddings. We use word embeddings of the languages projected to a common vector space on which a community-detection algorithm is applied to find clusters such that words that represent the same concept from different languages fall in the same group. We utilize these multilingual word clusters to perform query translation for Cross-Language Information Retrieval for three languages - English, Hindi and Bengali. We have experimented with the FIRE 2012 and Wikipedia datasets and have shown improvements over several standard methods like dictionary-based method, a transliteration-based model and Google Translate.
Code (0)
등록된 구현이 없습니다.
Tasks
Community DetectionInformation RetrievalMachine TranslationMultilingual Word EmbeddingsRetrievalTranslationTransliterationWord EmbeddingsSimilar Papers 제목 키워드 기반
Document Translation vs. Query Translation for Cross-Lingual Information Retrieval in the Medical Domain
We present a thorough comparison of two principal approaches to Cross-Lingual Information Retrieval: document translation (DT) and query translation (QT). Our experiments are conducted using the cross-lingual test collec…
Cross-Lingual Information RetrievalDocument TranslationInformation RetrievalMachine Translation+3Language Modelling with NMT Query Translation for Amharic-Arabic Cross-Language Information Retrieval
This paper describes our first experiment on Neural Machine Translation (NMT) based query translation for Amharic-Arabic Cross-Language Information Retrieval (CLIR) task to retrieve relevant documents from Amharic and Ar…
Information RetrievalLanguage ModelingLanguage ModellingMachine Translation+3Exploiting Neural Query Translation into Cross Lingual Information Retrieval
As a crucial role in cross-language information retrieval (CLIR), query translation has three main challenges: 1) the adequacy of translation; 2) the lack of in-domain parallel training data; and 3) the requisite of low …
Cross-Lingual Information RetrievalData AugmentationDomain AdaptationInformation Retrieval+4Learning to Weight Translations using Ordinal Linear Regression and Query-generated Training Data for Ad-hoc Retrieval with Long Queries
Ordinal regression which is known with learning to rank has long been used in information retrieval (IR). Learning to rank algorithms, have been tailored in document ranking, information filtering, and building large ali…
Document RankingInformation RetrievalLearning-To-Rankregression+2UsingWord Embeddings for Query Translation for Hindi to English Cross Language Information Retrieval
Cross-Language Information Retrieval (CLIR) has become an important problem to solve in the recent years due to the growth of content in multiple languages in the Web. One of the standard methods is to use query translat…
Information RetrievalRetrievalTranslationWord Embeddings+1