paper-with-me

Papers

Learning to Respond to Mixed-code Queries using Bilingual Word Embeddings

2019-06-01 · NAACL 2019 6 · Chia-Fang Ho, Jason Chang, Jhih-Jie Chen, Ching-Yu Yang

We present a method for learning bilingual word embeddings in order to support second language (L2) learners in finding recurring phrases and example sentences that match mixed-code queries (e.g., {``}接 受 sentence{''}) composed of words in both target language and native language (L1). In our approach, mixed-code queries are transformed into target language queries aimed at maximizing the probability of retrieving relevant target language phrases and sentences. The method involves converting a given parallel corpus into mixed-code data, generating word embeddings from mixed-code data, and expanding queries in target languages based on bilingual word embeddings. We present a prototype search engine, x.Linggle, that applies the method to a linguistic search engine for a parallel corpus. Preliminary evaluation on a list of common word-translation shows that the method performs reasonablly well.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceTranslationWord EmbeddingsWord Translation

Similar Papers 제목 키워드 기반

MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language Queries

2025-05-22 · Jonghwi Kim, Deokhyung Kang, Seonjeong Hwang, Yunsu Kim 외

Despite bilingual speakers frequently using mixed-language queries in web searches, Information Retrieval (IR) research on them remains scarce. To address this, we introduce MiLQ,Mixed-Language Query test set, the first …

BenchmarkingInformation RetrievalRetrieval

Word Embeddings for Code-Mixed Language Processing

2018-10-01 · EMNLP 2018 10 · Adithya Pratapa, Monojit Choudhury, Sunayana Sitaram

We compare three existing bilingual word embedding approaches, and a novel approach of training skip-grams on synthetic code-mixed text generated through linguistic models of code-mixing, on two tasks - sentiment analysi…

Machine TranslationPOSPOS TaggingSentiment Analysis+1

Deep Learning Techniques for Humor Detection in Hindi-English Code-Mixed Tweets

2019-06-01 · WS 2019 6 · Sushmitha Reddy Sane, Suraj Tripathi, Koushik Reddy Sane, Radhika Mamidi

We propose bilingual word embeddings based on word2vec and fastText models (CBOW and Skip-gram) to address the problem of Humor detection in Hindi-English code-mixed tweets in combination with deep learning architectures…

Deep LearningHumor DetectionWord Embeddings

hinglishNorm -- A Corpus of Hindi-English Code Mixed Sentences for Text Normalization

2020-10-18 · Piyush Makhija, Ankit Kumar, Anuj Gupta

We present hinglishNorm -- a human annotated corpus of Hindi-English code-mixed sentences for text normalization task. Each sentence in the corpus is aligned to its corresponding human annotated normalized form. To the b…

SentenceText NormalizationTranslation

hinglishNorm - A Corpus of Hindi-English Code Mixed Sentences for Text Normalization

2020-12-01 · COLING 2020 8 · Piyush Makhija, Ankit Kumar, Anuj Gupta

We present hinglishNorm - a human annotated corpus of Hindi-English code-mixed sentences for text normalization task. Each sentence in the corpus is aligned to its corresponding human annotated normalized form. To the be…

SentenceText NormalizationTranslation