Learning to Respond to Mixed-code Queries using Bilingual Word Embeddings
We present a method for learning bilingual word embeddings in order to support second language (L2) learners in finding recurring phrases and example sentences that match mixed-code queries (e.g., {``}接 受 sentence{''}) composed of words in both target language and native language (L1). In our approach, mixed-code queries are transformed into target language queries aimed at maximizing the probability of retrieving relevant target language phrases and sentences. The method involves converting a given parallel corpus into mixed-code data, generating word embeddings from mixed-code data, and expanding queries in target languages based on bilingual word embeddings. We present a prototype search engine, x.Linggle, that applies the method to a linguistic search engine for a parallel corpus. Preliminary evaluation on a list of common word-translation shows that the method performs reasonablly well.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceTranslationWord EmbeddingsWord TranslationSimilar Papers 제목 키워드 기반
MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language Queries
Despite bilingual speakers frequently using mixed-language queries in web searches, Information Retrieval (IR) research on them remains scarce. To address this, we introduce MiLQ,Mixed-Language Query test set, the first …
BenchmarkingInformation RetrievalRetrievalWord Embeddings for Code-Mixed Language Processing
We compare three existing bilingual word embedding approaches, and a novel approach of training skip-grams on synthetic code-mixed text generated through linguistic models of code-mixing, on two tasks - sentiment analysi…
Machine TranslationPOSPOS TaggingSentiment Analysis+1Deep Learning Techniques for Humor Detection in Hindi-English Code-Mixed Tweets
We propose bilingual word embeddings based on word2vec and fastText models (CBOW and Skip-gram) to address the problem of Humor detection in Hindi-English code-mixed tweets in combination with deep learning architectures…
Deep LearningHumor DetectionWord EmbeddingshinglishNorm -- A Corpus of Hindi-English Code Mixed Sentences for Text Normalization
We present hinglishNorm -- a human annotated corpus of Hindi-English code-mixed sentences for text normalization task. Each sentence in the corpus is aligned to its corresponding human annotated normalized form. To the b…
SentenceText NormalizationTranslationhinglishNorm - A Corpus of Hindi-English Code Mixed Sentences for Text Normalization
We present hinglishNorm - a human annotated corpus of Hindi-English code-mixed sentences for text normalization task. Each sentence in the corpus is aligned to its corresponding human annotated normalized form. To the be…
SentenceText NormalizationTranslation